{"id":"623579f6-0a61-47d6-8b87-12aa6ff40270","arxiv_id":"2501.04056","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive review of RNA secondary structure prediction methods, RNA modification detection tools, and the datasets and open problems connecting them.","lead":"This paper is a review of computational methods for predicting RNA secondary structure and detecting RNA chemical modifications. It surveys dozens of tools and databases and highlights the growing link between modifications such as m6A and RNA folding.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: review-level accuracy issues are minor and do not threaten the survey's central value.","rationale":"The paper is a review, so the load-bearing condition for its central claim is reliable representation of prior work rather than a new scientific result. I checked the major technical arcs: the thermodynamic, comparative, learning-based, and hybrid method descriptions align with the cited sources; the modification-prediction tool summaries are consistent with the methods they describe; and Section 4.3 accurately reports the Kierzek/Szabat m6A parameter sets and the ViennaRNA extension. The reader's weakest_assumption identifies the m6A parameter transfer to in vivo conditions, but that is a field-wide open question acknowledged implicitly through the review's 'challenges and opportunities' framing, not an internal error in the survey. The MODOMICS count discrepancy is real and concrete, but it is a localized data-resource error, not a flaw in the central argument that the review maps the field. The loose 'hybrid' categorization of RNA-FM and RNAErnie is a taxonomy quibble rather than a correctness failure. Since no substantial concern lands, the honest non-finding is appropriate, and the reader's UNVERDICTED verdict should remain unchanged.","tokens_in":29694,"tokens_out":10783,"duration_ms":103195,"concrete_test":"Audit Section 1 and Table 4 against the MODOMICS 2023 update (Cappannini et al., Nucleic Acids Research 2024) to confirm whether '335 natural RNA modifications' should be '335 natural modified residues,' and spot-check Table 2's 'hybrid' labels against the primary papers for RNA-FM and RNAErnie. If these are the only corrections needed, the UNVERDICTED verdict should stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this survey provides an accurate, structured map of RNA secondary-structure prediction methods, RNA modification prediction tools, relevant data sources, and their interplay. Because the paper offers no novel method or derivation, the correctness bar is faithful representation of the cited literature. On the substantive technical descriptions of the RSS predictors, modification predictors, datasets, and the m6A nearest-neighbor parameter discussion, I found no error that would mislead a method developer at the level the review claims. The MODOMICS count in Section 1 ('over 335 natural RNA modifications') conflicts with Table 4 ('more than 170 different RNA modifications, 429 modified residues, 335 natural ones'), and the 'hybrid' label for RNA-FM and RNAErnie in Table 2 is taxonomically loose. These are genuine blemishes, but they are localized and do not undermine the survey's core synthesis of methods, data, and opportunities.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a review of computational methods for RNA secondary structure (RSS) prediction and RNA modification prediction, together with the data resources used in these areas and the interplay between the two fields. It organizes RSS prediction into energy-based, comparative, learning-based, and hybrid approaches, surveys prediction tools for six RNA modification types, discusses the bidirectional effects of modifications and structure with emphasis on m6A and nearest-neighbor thermodynamic parameters, and lists relevant databases. The review closes with challenges such as data scarcity, pseudoknot handling, and therapeutic applications.","tokens_in":29851,"tokens_out":8573,"duration_ms":71761,"significance":"If accurate, this survey provides a useful structured map of a fast-moving field, and it is generally faithful to the cited literature. Its strengths include its breadth—covering method classes, tools, and data—and its focused discussion of m6A thermodynamics, including the Kierzek et al. and Szabat et al. parameter sets and the ViennaRNA modified-base support. It also explicitly flags data biases such as redundancy in ArchiveII and limitations of PDB-derived RNA sets. The review does not introduce new methods, so its value rests on the reliability of its synthesis; the issues found below are local and do not undermine the overall contribution.","major_comments":[],"minor_comments":[{"comment":"Section 1 states that MODOMICS catalogs 'over 335 natural RNA modifications', but Section 5.2 and Table 4 report 'more than 170 different RNA modifications, 429 different RNA modified residues (335 natural ones)'; this internal inconsistency is a factual error and should be corrected.","section":"Sec. 1 vs. Sec. 5.2/Table 4"},{"comment":"RNA-FM and RNAErnie are labeled as hybrid (H) in Table 2 and discussed under hybrid methods in Section 2.4, yet they are pre-trained language models rather than combinations of thermodynamic, comparative, and learning strategies as defined in Table 1; please clarify the hybrid definition or reclassify these entries.","section":"Table 2 / Sec. 2.4"},{"comment":"The subsection title 'RSS related RNA modification datasets' is misleading because ENCORI and eCLIP contain protein-RNA interaction data rather than RNA modification data; consider renaming the subsection or explicitly explaining their indirect relevance.","section":"Sec. 5.2"},{"comment":"The TORNADO entry lists 'Single sequence (FASTA, Stockholm)' as input, but Stockholm is a multiple-sequence alignment format; please clarify what input TORNADO actually accepts.","section":"Table 2"},{"comment":"There are numerous typographical and formatting errors, including 'Waston-Crick', 'Felsenstain', 'unaired region' (should be 'unpaired'), 'the the NIH grants', 'seqeunces', 'Lewi et al.' (should be 'Lewis et al.'), 'Tanzeret al.', 'Conventional' (should be 'Conventionally'), and inconsistent spacing/casing such as 'L INEAR FOLD', 'RNAFOLD', and 'MODOMICs'.","section":"Throughout"},{"comment":"Several of the cited modification prediction tools (e.g., H2Opred, Meta-2OM, ac4C-AFL, MST-m6A) are authored by a co-author of this review; adding a disclosure statement would improve transparency.","section":"Refs / disclosure"}],"recommendation":"minor_revision","confidential_remarks":"The paper is within scope for q-bio.BM. The self-citation pattern for tools authored by one of the co-authors is worth editorial awareness, but it does not affect the technical content. The MODOMICS count inconsistency and the hybrid-label issue are local and should be fixed in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a review, and a fairly good one. It doesn't introduce new methods, data, or results — the authors say so explicitly — but it does something useful: it maps the methodological arc of RSS prediction from Nussinov/Zuker through comparative SCFGs to deep learning and hybrid foundation models, then connects that to computational RNA modification prediction and the specific question of how m6A and other modifications change folding energetics. The m6A nearest-neighbor parameter story (Kierzek, Szabat, RNAstructure and ViennaRNA integration) is the most valuable part of the review; it is technically accurate and gives method developers a concrete entry point.\n\nThe tool tables (Tables 2 and 3) are genuinely handy. I spot-checked descriptions of RNAfold, RNAstructure, LinearFold, CONTRAfold, SPOT-RNA, MXfold2, RNA-FM, and the modification predictors; the characterizations match the primary sources. The survey is also honest about limitations, e.g., the bias in PDB and redundancy in ArchiveII.\n\nSoft spots are minor and local. The MODOMICS count is inconsistent: Section 1 says 'over 335 natural RNA modifications,' Table 4 says 'more than 170 different RNA modifications, 429 different RNA modified residues (335 natural ones).' One of those is wrong or the wording is sloppy. Also, Table 2 classifies RNA-FM and RNAErnie as 'hybrid,' which is loose — they are pre-trained foundation models fine-tuned for RSS; 'hybrid' is usually reserved for mixing thermodynamic and learning. There are typos and citation-format slips throughout (e.g., 'seqeunces', 'Waston-Crick', the repeated 'the the'). None of these undercut the central synthesis.\n\nThe weakest part is Section 4, which is more suggestive than rigorous: the authors assert interplay between modifications and structure mostly by citing reviews and a few primary papers, and the reader should not expect a critical evaluation of benchmark quality. The review takes ArchiveII, bpRNA, RNAStralign, and the m6A parameters at face value. That's acceptable for a survey but worth saying out loud.\n\nWho should read it: newcomers to RNA bioinformatics, method developers wanting a quick map of tool classes and datasets, and anyone working on modification-aware folding. It deserves a serious referee — an editor should send it out, though the referee should ask for the MODOMICS count to be fixed and the hybrid taxonomy tightened. I'd cite it and bring it to a reading group.","headline":"A competent, current survey of RNA secondary structure prediction and modification tools; no new methods, but the synthesis and m6A parameter discussion make it worth refereeing.","tokens_in":30388,"tokens_out":1944,"would_cite":true,"duration_ms":19414,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"RNA secondary structure and RNA modifications are one coupled system, and prediction methods are converging on modeling them together.","keywords":["RNA secondary structure prediction","RNA modifications","m6A","nearest-neighbor thermodynamics","deep learning","RNA foundation models","epitranscriptomics","benchmark datasets"],"falsifier":"Run a head-to-head test of m6A-aware folding versus standard folding on in-cell RNA structure probing data such as SHAPE-MaP maps of m6A-containing transcripts; if the m6A parameters do not improve accuracy over unmodified parameters, the review's core recommendation for modified-RNA prediction collapses. Alternatively, retrain a top deep-learning model on a redundancy-filtered holdout from ArchiveII or bpRNA-1m and show that method rankings invert.","tokens_in":29512,"feed_emoji":"🧬","tokens_out":5527,"duration_ms":47944,"temperature":0.7,"pith_summary":"The paper argues that RNA secondary structure prediction and RNA modification identification have grown from separate toolboxes into a single intertwined problem: chemical modifications such as m6A change the free-energy landscape of folding, while structural context determines where modifications are installed. It organizes four decades of prediction methods into thermodynamic, comparative, learning-based, and hybrid families, then catalogs modification-prediction tools by modification type. Its distinctive contribution is to make explicit the two-way link between structure and modification, and to survey the data resources and energy parameters, such as m6A-expanded nearest-neighbor parameters, that now make joint modeling possible. If the review's reading of the field is right, the near-term payoff is structure prediction that accounts for modifications and modification prediction that uses structure features, improving RNA therapeutics design.","feed_headline":"RNA structure and modifications must be modeled together","feed_subtitle":"A survey of four method families shows how m6A energy parameters and deep learning make joint prediction practical.","key_machinery":"The load-bearing object is the nearest-neighbor free energy model: a set of empirically measured parameters that decomposes an RNA secondary structure into loop and stacking units whose free energies add. The review shows how this same machinery, originally built on unmodified nucleotides, is being extended to modified RNAs, most completely for m6A, with full parameter sets from optical melting experiments, and how the Zuker and McCaskill dynamic programming algorithms use these parameters to predict modified structures. On the modification side, the counterpart machinery is the sequence-feature classifiers, including SVM, CNN, transformer, and ensemble methods, that consume RNA structural encodings such as RNAfold MFE values to predict modification sites. The review's organizational machinery is its four-way classification of RNA secondary structure prediction methods.","core_discovery":"The central discovery the paper seeks to establish is that RNA secondary structure and RNA modifications are not separate phenomena but a coupled regulatory system: RNA secondary structure motifs guide where writers like METTL3/METTL14 install modifications, and modifications such as m6A, A-to-I editing, and pseudouridine reshape the structure by changing base-pair stabilities and annealing kinetics. The review supports this by tracing method progression from Nussinov-style dynamic programming and Turner nearest-neighbor free energy models, through comparative SCFG methods, to deep learning and foundational RNA language models, and by highlighting the emergence of energy parameters for modified nucleotides, particularly the complete m6A nearest-neighbor parameter set measured by optical melting experiments and integrated into RNAstructure and ViennaRNA. The intended conclusion is that accurate prediction of functional RNA behavior requires models that fold modified sequences with modification-aware energy parameters and that exploit structural context to predict modification sites.","pith_inferences":["A testable extension the paper does not run: benchmark m6A-aware folding against transcriptome-wide in-cell probing data such as SHAPE-MaP to see whether melting-derived parameters improve accuracy over unmodified parameters in vivo.","The review's taxonomy implies that energy-based and learning-based approaches are converging; a plausible next step is fully differentiable folding that trains end-to-end on modified RNA structure data.","If the m6A-switch mechanism generalizes to other modifications, structure-aware modification prediction could become a general epitranscriptomics design rule rather than a per-modification special case."],"forward_implications":["Structure prediction tools that ignore modifications will systematically mispredict folding of modified transcripts, especially in m6A-rich regions such as 3' UTRs.","Including m6A-aware nearest-neighbor parameters in mainstream packages like RNAstructure and ViennaRNA makes modified-RNA folding computationally practical today.","Using RNA secondary structure features in modification predictors improves site prediction for m6A, 2'-O-methylation, and other marks.","Deep learning and RNA foundation models pre-trained on large sequence collections are the likely route to overcome data scarcity and bias in RNA structure prediction.","Joint modeling of structure and modification can inform the rational design of RNA vaccines and gene therapies, for example with N1-methylpseudouridine-modified messages."],"supporting_citations":[{"why":"Supplies the Turner nearest-neighbor free-energy parameters used by most energy-based folding methods reviewed.","marker":"51"},{"why":"Provides the complete set of nearest-neighbor parameters for m6A-modified RNA and shows their use in RNAstructure.","marker":"18"},{"why":"Refines the m6A parameters through roughly 100 optical melting experiments and validates them against measured folding stabilities.","marker":"104"},{"why":"Extends the ViennaRNA package with energy parameters and constraints for six modified RNA bases.","marker":"105"},{"why":"Establishes the m6A-switch mechanism linking m6A-dependent structural remodeling to RNA-protein interactions.","marker":"101"},{"why":"Represents the deep-learning approach to RSS prediction, including noncanonical and pseudoknotted pairs.","marker":"12"},{"why":"Establishes the RNA foundation model approach pre-trained on unannotated sequences for structure and function prediction.","marker":"14"},{"why":"Provides the ArchiveII benchmark dataset widely used to train and test RSS prediction methods.","marker":"114"},{"why":"Catalogs RNA modifications and modified residues, underpinning the modification section's census.","marker":"17"}],"fun_headline_variants":["RNA structure and modification code are inseparable","Joint prediction of RNA folds and marks: the path forward","m6A energy parameters unlock coupled RNA folding","Stop modeling RNA structure without modifications","RNA secondary structure and epitranscriptomics: one system"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's guidance assumes that thermodynamic parameters measured on short synthetic RNAs in the laboratory, especially the m6A nearest-neighbor values, transfer faithfully to how modified RNAs fold inside cells, and that the benchmark datasets it highlights are representative enough to rank methods reliably.","fun_headline_variants_meta":{"raw":{"variants":["RNA structure and modification code are inseparable","Joint prediction of RNA folds and marks: the path forward","m6A energy parameters unlock coupled RNA folding","Stop modeling RNA structure without modifications","RNA secondary structure and epitranscriptomics: one system"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1245,"prompt_tokens":902,"completion_tokens":343,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":273}},"tokens_in":518,"tokens_out":343,"duration_ms":3992,"temperature":1.0,"reasoning_tokens":273,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:51:50.514754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a head-to-head test of m6A-aware folding versus standard folding on in-cell RNA structure probing data such as SHAPE-MaP maps of m6A-containing transcripts; if the m6A parameters do not improve accuracy over unmodified parameters, the review's core recommendation for modified-RNA prediction collapses. Alternatively, retrain a top deep-learning model on a redundancy-filtered holdout from ArchiveII or bpRNA-1m and show that method rankings invert.","supporting_citations":[{"cited_title":"A test and refinement of folding free energy nearest neighbor parameters for RNA including N6-methyladenosine","cited_arxiv_id":null,"evidence_quote":"Refines the m6A parameters through roughly 100 optical melting experiments and validates them against measured folding stabilities."},{"cited_title":"Modified RNAs and predictions with the ViennaRNA Package","cited_arxiv_id":null,"evidence_quote":"Extends the ViennaRNA package with energy parameters and constraints for six modified RNA bases."},{"cited_title":"N 6-methyladenosine-dependent RNA structural switches regulate RNA– protein interactions","cited_arxiv_id":null,"evidence_quote":"Establishes the m6A-switch mechanism linking m6A-dependent structural remodeling to RNA-protein interactions."},{"cited_title":"Exact calculation of loop formation probability identifies folding motifs in RNA secondary structures","cited_arxiv_id":null,"evidence_quote":"Provides the ArchiveII benchmark dataset widely used to train and test RSS prediction methods."}],"review_version":1}