{"id":"c7cba197-e867-4a68-8edc-3a3ead0b7c41","arxiv_id":"2504.17162","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"This arXiv paper is a narrative review that organizes AI/ML methods for RNA subcellular localization prediction into sequence-based, image-based, and hybrid categories.","lead":"This paper reviews machine learning methods that predict where RNAs live inside cells, using sequence data, microscopy images, or both. It is a survey, not a new method, and its value depends on whether its citations and summaries can be trusted.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Citation-number errors in Sections III.B, III.C, and IV send readers to unrelated papers, breaking the review's core function as a reliable map of the literature.","rationale":"The reader's verdict is REJECT, and the stress test confirms that the most load-bearing weakness is citation integrity. The reader's weakest_assumption identified exactly this: each bracketed citation must point to a paper supporting the attached claim. My independent check of the manuscript and reference list confirms concrete mismatches: [14] is Jeffery et al. 1983 rather than the Bento toolkit; [188] is Kiskowski et al. 2009 rather than a multimodal deep learning paper; [194] is Huang et al. 2019 rather than Savulescu et al. 2021. These are mechanical but material defects in a review paper, where citations are the evidence trail. The unresolved 'because xxx' placeholder in Section II.B further indicates the draft is unfinished. I found no reason to disagree with the reader's verdict, and no additional adversarial concern is needed. The broad taxonomy and challenges sections have genuine value, but the current manuscript would mislead any reader who follows its citations, so the rejection stands unchanged. If the citation errors and placeholder are corrected, the paper could be reconsidered, but that is a revision decision rather than a change to the validity of the current verdict.","tokens_in":30541,"tokens_out":2330,"duration_ms":22832,"concrete_test":"Perform a systematic citation-claim audit: for every narrative sentence containing bracketed references, compare the sentence's claim against the title and abstract of each cited reference. Specifically verify whether (1) reference [14] is the Bento/Mah et al. random-forest image-based method, (2) reference [188] supports the multimodal-data claim about histopathology images in Section IV, and (3) references [188][194] correspond to the hybrid frameworks described in Section III.C. If any of these fail, the manuscript must be revised with corrected citation numbers before it can serve as a reliable review.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it is a comprehensive, reliable review of AI-based RNA subcellular localization prediction. In a review, citations are the primary evidence: each bracketed reference must point to a paper that supports the sentence it is attached to. This assumption fails in multiple directly verifiable places. Section III.B states 'In [14], Clarence et al. applied RF to mRNA localization as a multi-label classification problem.' Reference [14] is Jeffery et al. 1983, 'Localization of actin messenger RNA during early ascidian development' — a wet-lab developmental biology paper, not a random-forest image classifier. The actual Bento toolkit by Mah et al. is correctly listed in Table 3 as reference [179], so the narrative citation is shifted or wrong. Section III.C cites hybrid frameworks as [188][194], while the described frameworks are [178] (Wang et al.) and [180] (Savulescu et al.); [188] is Kiskowski et al. 2009 on Ripley's K-function, and [194] is Huang et al. 2019 on breast cancer survival analysis. Section IV attributes to 'Wang et al. [188]' a claim about genome profile data lacking paired histopathology images, but [188] is again Kiskowski et al. 2009. These are not interpretive disagreements; the bibliographic titles and DOIs unambiguously show the cited papers do not support the claims. Additionally, Section II.B contains the unresolved placeholder 'deep neural networks with the attention mechanism were usually preferred because xxx', indicating an unfinished sentence. For a narrative review, citation integrity is the load-bearing condition, and these errors directly undermine the stated purpose of serving as a valuable resource for researchers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey-style review of artificial intelligence and machine learning methods for predicting the subcellular localization of RNAs, including mRNA, lncRNA, and miRNA. The paper organizes methods into three categories: sequence-based, image-based, and hybrid approaches that combine both. It presents two large tables summarizing sequence-based and image-based/hybrid predictors, discusses feature extraction and classification algorithms, and closes with a section on challenges and future directions, including data scarcity, over-reliance on sequence features, and the lack of standardized benchmarks.","tokens_in":30740,"tokens_out":3227,"duration_ms":28822,"significance":"If the review were accurate and complete, it would be a useful entry point for researchers entering the field, because the taxonomy of sequence-, image-, and hybrid-based methods is coherent and the chronological tables cover many recent publications. The manuscript also explicitly identifies open problems such as small dataset sizes, class imbalance, and missing benchmark datasets. However, the central claim of being a reliable, comprehensive map of the literature is undermined by directly verifiable citation-to-claim mismatches in multiple sections, including an unresolved placeholder text in the main body. Since a review's primary value is the trustworthiness of its references, these errors are load-bearing rather than cosmetic.","major_comments":[{"comment":"The sentence 'In [14], Clarence et al. applied RF to mRNA localization as a multi-label classification problem' cites reference [14], which is Jeffery et al. 1983, 'Localization of actin messenger RNA during early ascidian development' (Dev. Biol. 99:408-417). That is a wet-lab developmental biology paper, not a random-forest image classifier. The described method is the Bento toolkit, which is correctly listed in Table 3 as reference [179]. This mis-citation sends the reader to an unrelated primary source and directly contradicts the review's function as a reliable pointer to the literature.","section":"III.B, paragraph on image-based algorithms"},{"comment":"The text states that 'hybrid algorithms like multimodal fusion models [188] and layered neural networks [194] were only proposed but not implemented.' Reference [188] is Kiskowski et al. 2009, a paper on Ripley's K-function for analyzing domain size, and reference [194] is Huang et al. 2019, a breast cancer survival analysis method (SALMON). The actual hybrid frameworks described elsewhere in the paper are [178] (Wang et al. 2023) and [180] (Savulescu et al. 2021). The bracketed citations in this sentence are therefore entirely disconnected from the claims they are supposed to support.","section":"III.C, paragraph on hybrid methods"},{"comment":"The sentence 'as Wang et al. [188] points out, the genome profile data may lack paired histopathology image data for multimodal deep learning approaches to localization prediction' again cites [188], which is Kiskowski et al. 2009 on Ripley's K-function. That paper contains no such claim about genome profile data or histopathology images. This is the same mis-citation appearing in a different context, demonstrating that the problem is not isolated to one sentence but pervades the manuscript's use of references.","section":"IV, first paragraph of challenges"},{"comment":"The text 'Liu et al. in a 2014 study [113] as a hybrid approach that uses both XGBoost and convolutional neural networks' cites reference [113], which is Liu et al. 2014, 'iDNA-Prot|dis: Identifying DNA-Binding Proteins by Incorporating Amino Acid Distance-Pairs and Reduced Alphabet Profile into the General Pseudo Amino Acid Composition.' That work is about DNA-binding protein prediction and does not use XGBoost or convolutional neural networks. This is another example of a citation that does not support the attached claim.","section":"II.B, paragraph on ensemble methods"},{"comment":"The sentence 'In many cases [136][142][147][151], deep neural networks with the attention mechanism were usually preferred because xxx' contains an unresolved placeholder 'xxx' and is not a complete sentence. This is not a minor typo; it is an incomplete fragment in the main text that signals the manuscript has not been carefully finalized. The surrounding discussion of attention mechanisms is also not supported by the specific citation numbers, which refer to a mix of papers that do not all focus on attention-based deep networks.","section":"II.B, paragraph on deep learning methods"}],"minor_comments":[{"comment":"The phrase 'RNA s' in the first sentence contains an extra space before the plural 's'; the same pattern appears elsewhere and should be corrected throughout.","section":"Abstract"},{"comment":"The caption states that nuclear envelope 3D labels are used to discriminate distributions 'that are undisguisable' from the 2D view; the intended word is likely 'indistinguishable.'","section":"Table 3 caption"},{"comment":"The phrase 'Ripley's L-functions [188]' is imprecise because [188] specifically discusses the K-function and its derivatives; the text should name the exact quantity to match the cited source.","section":"III.A, paragraph on image-based features"}],"recommendation":"reject","confidential_remarks":"The manuscript is a review whose acceptance hinges on the accuracy of its citation-to-claim correspondences. The verified mismatches in Sections III.B, III.C, IV, and II.B are not isolated: they appear in the core sections that describe the state of the art. An unresolved placeholder ('because xxx') in the main text reinforces the impression that this draft has not been through a careful verification pass. I would not consider acceptance or even a routine revision; the authors would need to re-audit every referenced claim against its cited source, a process that goes well beyond local fixes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the review gives a useful map of a messy field, but the map has wrong labels — several citations point to unrelated papers and one sentence is unfinished. Do not let anyone use it as-is.\n\nThe paper organizes RNA subcellular localization prediction into sequence-based, image-based, and hybrid methods, with good tables of predictors and a sensible taxonomy. The introduction covers the biology accurately, and the challenges section (data scarcity, over-reliance on sequence features, missing benchmarks) names real bottlenecks. For a reader new to the area, the structure alone is helpful.\n\nThe problems are in the citation integrity. In Section III.B, the text says 'In [14], Clarence et al. applied RF to mRNA localization.' Reference [14] is Jeffery et al. 1983, a wet-lab paper on ascidian embryos. In Section III.C, hybrid frameworks are cited as [188][194], but [188] is Kiskowski et al. on Ripley's K-function and [194] is Huang et al. on breast cancer survival. In Section IV, the claim about lacking paired histopathology images is attributed to 'Wang et al. [188]' — again Kiskowski. These are not interpretive disagreements; the titles and DOIs are unambiguous. In a review, the bibliography is the evidence, and these errors send readers to papers that do not support the sentences. There is also an unresolved placeholder in Section II.B — 'deep neural networks with the attention mechanism were usually preferred because xxx.' That is an unfinished draft, not a stylistic quibble.\n\nThe paper is honest that hybrid methods are proposed but not yet implemented, which is fine, but the surrounding text twice suggests they improve prediction without citing an implemented system. That overclaims, though mildly.\n\nWho is this for? A graduate student wanting a bird's-eye view of existing predictors would get value from the tables and the three-way categorization. But they would be misled by the citations in their current state.\n\nRecommendation: reject this version. The taxonomy and tables are worth salvaging; the authors should re-check every reference against the sentence it supports, fill the placeholder, and tone down the hybrid-method enthusiasm. I would support sending a revised version to peer review. If the editor is deciding now, I would not desk-reject outright — the paper's scope is right and a referee can give the authors a concrete fix list — but the manuscript is not publishable in this form.","headline":"A useful but unfinished survey: the taxonomy and tables are solid, yet multiple citations point to unrelated papers, so the review cannot be trusted in its current form.","tokens_in":31362,"tokens_out":2613,"would_cite":false,"duration_ms":25262,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review organizes AI-based RNA subcellular localization prediction into sequence-based, image-based, and hybrid methods, and argues hybrid multimodal fusion is the field's most promising next step.","keywords":["RNA subcellular localization","long non-coding RNA","mRNA localization","microRNA localization","sequence-based prediction","image-based prediction","hybrid multimodal methods","deep learning"],"falsifier":"Take a random sample of entries from the review's method tables, retrieve each cited paper, and verify that it describes the algorithm, features, and localization set listed; any substantial mismatch, or any established method absent after a systematic search of the same literature, would show the review is not a reliable map.","tokens_in":30278,"feed_emoji":"🧬","tokens_out":6557,"duration_ms":58841,"temperature":0.7,"pith_summary":"RNA's job depends on where it sits in the cell, but wet-lab localization is slow and costly. This review surveys AI and machine-learning predictors that instead read RNA sequence, microscopy images, or both, and sorts them into three families: sequence-based, image-based, and hybrid. It argues the sequence-based family is mature and dominant, image-based methods add spatial information but are underused, and hybrid fusion methods are the most promising direction even though none has been implemented yet. The review also identifies data scarcity, imbalanced labels, and missing benchmark datasets as the obstacles that currently cap prediction quality. If it is right, a newcomer can use its tables as a working map of the field and concentrate effort on multimodal fusion rather than another sequence-only variant.","feed_headline":"Three AI routes predict where RNAs live in cells","feed_subtitle":"Sequence, image, and hybrid methods compared; fusion is promising but not yet built.","key_machinery":"The organizing device is a three-way taxonomy of methods by input data type: sequence-based, image-based, and hybrid. The taxonomy does the argument's work by turning dozens of separate predictors into three families, each with its own typical features (k-mers, PseKNC, one-hot, physicochemical values for sequences; spatial statistics and raw image tensors for images), algorithms, and failure modes. The companion pipeline—data preparation, feature extraction, feature selection, classification, and single- or multi-label assignment—is the standard template against which each method is described and compared.","core_discovery":"The paper's central claim is that the computational prediction of RNA subcellular localization has converged on a shared pipeline—input data, feature extraction, feature selection, model training, and localization assignment—and that the methods filling that pipeline divide cleanly by input modality. Sequence-based methods, the largest group, rely on k-mer composition, pseudo nucleotide composition, physicochemical features, one-hot encoding, and embeddings, and use algorithms from support vector machines and random forests to CNNs, LSTMs, graph networks, and transformers. Image-based methods extract spatial statistics such as point distributions, Ripley's L-functions, and morphological features from smFISH-style images, or feed preprocessed images directly to a CNN. Hybrid methods that fuse both modalities are described as the natural synthesis, but the paper states that at the time of writing such algorithms had been proposed yet not implemented. The review's own conclusion is that the field's progress is currently limited more by data—small, imbalanced, single-RNA-type datasets and the absence of standardized benchmarks—than by model architecture.","pith_inferences":["If hybrid fusion matures, its first decisive test will be whether adding image features to sequence models beats the best sequence-only baselines on the same benchmark; the review does not perform that comparison, so the claim remains a prediction.","The review's taxonomy suggests a concrete research agenda: generate more synthetic smFISH data and automated annotations to relieve the image-data bottleneck, then reuse the well-established sequence encoders as the sequence branch of a fused model.","The emphasis on interpretability in recent deep-learning predictors implies that future methods may be judged not only by accuracy but by whether their attention or gradient scores identify the same cis-regulatory motifs that wet-lab experiments localize."],"forward_implications":["A researcher choosing a predictor for mRNA or lncRNA can use the review's method tables as a baseline list; any new sequence-based model should be compared against the families represented there.","Multi-label prediction is the realistic framing, since RNAs often co-localize or shuttle between compartments; single-label results should be read with that limitation in mind.","Image-based approaches are complementary to sequence methods but currently depend on simulated data and costly annotation, so their growth is tied to automated image analysis.","Hybrid multimodal fusion is predicted to be the field's next breakthrough; the review's central gap is that no such system had been published at the time of writing.","Standardized large-scale benchmarks, built from resources like RNALocate and extended beyond fewer-than-ten-thousand mRNA samples, are a necessary condition for reliable comparison."],"supporting_citations":[{"why":"Grounds the sequence-based category as one of the first mRNA localization predictors using k-mer features and deep recurrent networks.","marker":"[112]"},{"why":"Supplies the multi-label mRNA prediction architecture and the AUC range (0.4–0.8) the review cites for state-of-the-art methods.","marker":"[131]"},{"why":"Establishes the lncLocator stacked ensemble classifier, a reference point for single-label lncRNA sequence predictors.","marker":"[159]"},{"why":"Supports the challenge claim that k-mer composition accounts for about 90% of decision-making in a widely used sequence-only model.","marker":"[158]"},{"why":"Provides the RNALocate resource that anchors the discussion of data availability and benchmark limitations.","marker":"[95]"},{"why":"Grounds the image-based category with a SqueezeNet CNN classifying mRNA localization patterns from simulated smFISH images.","marker":"[181]"},{"why":"Supplies the image-feature framework using clustering and random forests with Ripley's L-functions for RNA localization.","marker":"[182]"},{"why":"Supports the image-feature discussion through the Bento toolkit's spatial features for point distributions in spatial transcriptomics.","marker":"[179]"},{"why":"Represent the hybrid direction the review identifies as promising, proposed as heterogeneous multimodal learning but not yet implemented.","marker":"[180]"}],"fun_headline_variants":["RNA localization AI: three routes, no hybrid yet","Data, not algorithms, holds back RNA localization AI","Missing fusion: the gap in RNA localization prediction","Sequence vs image: RNA localization methods reviewed","Hybrid models promised, not yet built for RNA location"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's usefulness as a map rests on every bracketed citation pointing to a paper that actually contains the method or result attributed to it; if citation-to-claim matches are wrong, readers can be misled even when the three-way taxonomy is correct.","fun_headline_variants_meta":{"raw":{"variants":["RNA localization AI: three routes, no hybrid yet","Data, not algorithms, holds back RNA localization AI","Missing fusion: the gap in RNA localization prediction","Sequence vs image: RNA localization methods reviewed","Hybrid models promised, not yet built for RNA location"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0004,"raw_usage":{"total_tokens":2121,"prompt_tokens":1006,"completion_tokens":1115,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":1041}},"tokens_in":622,"tokens_out":1115,"duration_ms":8118,"temperature":1.0,"reasoning_tokens":1041,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:47:19.335413+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of entries from the review's method tables, retrieve each cited paper, and verify that it describes the algorithm, features, and localization set listed; any substantial mismatch, or any established method absent after a systematic search of the same literature, would show the review is not a reliable map.","supporting_citations":[{"cited_title":"Prediction of mRNA subcellular localization using deep recurrent neural networks,","cited_arxiv_id":null,"evidence_quote":"Grounds the sequence-based category as one of the first mRNA localization predictors using k-mer features and deep recurrent networks."},{"cited_title":"M3Loc: multi-label mRNA subcellular localization prediction and analysis based on multi-head self-attention mechanism,","cited_arxiv_id":null,"evidence_quote":"Supplies the multi-label mRNA prediction architecture and the AUC range (0.4–0.8) the review cites for state-of-the-art methods."},{"cited_title":"The lncLocator: a subcellular localization predictor for long non-coding RNAs based on a stacked ensemble classifier,","cited_arxiv_id":null,"evidence_quote":"Establishes the lncLocator stacked ensemble classifier, a reference point for single-label lncRNA sequence predictors."},{"cited_title":"Prediction of LncRNA Subcellular Localization with Deep Learning from Sequence Features,","cited_arxiv_id":null,"evidence_quote":"Supports the challenge claim that k-mer composition accounts for about 90% of decision-making in a widely used sequence-only model."},{"cited_title":"RNALocate v2.0: an updated resource for RNA subcellular localization with increased coverage and annotation,","cited_arxiv_id":null,"evidence_quote":"Provides the RNALocate resource that anchors the discussion of data availability and benchmark limitations."},{"cited_title":"Bento: a toolkit for subcellular analysis of spatial transcriptomics data,","cited_arxiv_id":null,"evidence_quote":"Supports the image-feature discussion through the Bento toolkit's spatial features for point distributions in spatial transcriptomics."}],"review_version":1}