{"id":"d0da8e23-142e-4cb7-b463-d2f4adcf7ad6","arxiv_id":"2411.16084","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A PRISMA-based scoping review finds 26 studies (2021-2024) applying NLP and transformers to genomics, with k-mer tokenization and BERT-style models dominating.","lead":"This paper is a structured review of 26 studies that use natural language processing and large language models to analyze DNA and RNA sequences. It summarizes common tokenization methods, transformer models, and prediction tasks such as transcription factor binding and methylation site detection.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 26-study corpus is not reproducible from the reported search protocol: Appendix B contains malformed queries and Section 2.2 lists inconsistent sources, so the review's landscape claim rests on an unverified completeness assumption.","rationale":"The reader's weakest_assumption identified search completeness and reliability of reported metrics. My independent read converges on the same area, with a sharper focus: the manuscript's own Appendix B makes the search protocol non-reproducible. The ACM query as printed has syntactically invalid segments such as '[Abstract: or]' and a trailing 'or'; the Scopus query similarly ends with an operator; the WoS query has a dangling 'OR'; and the body text lists additional databases that are absent from the abstract and appendix. The PRISMA flow chart is referenced as Figure 1 but not included in the text, so the screening arithmetic cannot be checked. Because the review's headline finding is a characterization of the current NLP-genomics landscape, the completeness of the underlying corpus is the load-bearing assumption. If the corrected search retrieves materially more studies, the directional claims about tokenization and transformer models could still be true but would be insufficiently supported by this review. The bibliographic error involving reference [5] (same title as reference [43] but different generic authors) reinforces that source verification is the weak point, but it does not by itself refute the central claim. This read keeps the reader's conditional verdict: the paper should not be relied on as a definitive map until the search is re-run and the included corpus is confirmed. I do not recommend acceptance or rejection on the current evidence, only that the stated conditions be verified.","tokens_in":15453,"tokens_out":9028,"duration_ms":82840,"concrete_test":"Have an independent searcher execute each Appendix B query exactly as printed in the named databases and record the number of hits. Then execute corrected versions: fix the ACM bare-OR and trailing-operator errors, remove dangling boolean terms from the Scopus/WoS/ACM strings, and reconcile the database list in Section 2.2. Compare the set of eligible records from the corrected search with the 26 included studies. If the corrected search retrieves additional eligible studies, or if the printed queries return zero results or error messages, the review's completeness claim fails and the synthesis must be revised or re-scoped.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of this scoping review is descriptive: 26 selected studies show tokenization and transformer models enhancing genomic processing, especially for regulatory annotation prediction. For that claim to land, the 26 studies must be the relevant population, not an arbitrary sample. That condition is not currently met. Section 2.2 names different databases in the abstract (PubMed, Medline, Scopus, Web of Science, Embase, ACM) and in the body (adding IEEE Xplore, Google Scholar, Semantic Scholar, Ovid MEDLINE, Ovid EMBASE), so the actual search scope is ambiguous. More importantly, Appendix B prints queries that cannot have been executed as written: the ACM query contains bare '[Abstract: or]' and '[Abstract: or \"nlp\" or \"llm\" or]' fields, and several strings end with dangling boolean operators. A PRISMA-style review that cannot reproduce its record-identification step cannot support a definitive landscape claim. Bibliographic verification is also weak: reference [5] carries the same title as the legitimate reference [43] but is attributed to different, generic authors, which is a checkable error that should be corrected. The reader's assumption about search completeness is therefore the load-bearing one; metric accuracy is secondary because even accurate metrics from an incomplete corpus would not validate the synthesis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a PRISMA-style scoping review of NLP and large language model (LLM) techniques applied to genomic sequencing data. The authors report a systematic search of bibliographic databases, identify 26 studies published between 2021 and April 2024, and synthesize them along three axes: tokenization (k-mers, BPE, fixed nucleotide tokenization), transformer architectures (BERT variants, encoder blocks, attention mechanisms), and downstream tasks centered on regulatory annotation prediction (methylation sites, transcription-factor binding, promoter prediction, RNA interactions, cancer-related tasks). The review concludes that tokenization and transformer models enhance genomic data processing, particularly for regulatory annotation prediction, and discusses challenges such as model interpretability, tokenizer selection, and computational resource demands.","tokens_in":15662,"tokens_out":4154,"duration_ms":38580,"significance":"If the search strategy were reproducible and the bibliographic record reliable, this review would be a useful structured map of a fast-moving field, consolidating model names, tasks, data types, and reported performance metrics in one place. The paper's strengths are its explicit PRISMA framing, the detailed per-study table (Table 1), the summary metric table (Table 2), and an unusually candid limitations section. However, the central descriptive claim—that these 26 studies represent the relevant population of NLP-for-genomics work—currently rests on a search protocol that cannot be reproduced from the manuscript, and there are verifiable bibliographic and factual errors. These are correctable in revision but are load-bearing for a scoping review's principal purpose.","major_comments":[{"comment":"The information sources are described inconsistently: the abstract lists PubMed, Medline, Scopus, Web of Science, Embase, and ACM, while §2.2 adds IEEE Xplore, Google Scholar, Semantic Scholar, and Ovid MEDLINE/EMBASE. More seriously, the search strategies in Appendix B cannot have been executed as printed: the ACM query contains bare bracket fields such as '[Abstract: or]' and '[Abstract: or \"nlp\" or \"llm\" or]' and ends with dangling Boolean operators, and no search strings are given for IEEE Xplore, Google Scholar, or Semantic Scholar. Because the review's landscape claim depends on identifying the relevant study population, the reported protocol does not currently support the statement that the 26 included studies constitute a complete or representative corpus. Please provide corrected queries and report per-database hit counts and deduplication numbers.","section":"§2.2 and Appendix B"},{"comment":"References [5] and [43] carry the same title, 'Effect of tokenization on transformers for biological sequences,' but [5] is attributed to 'John Smith, Anne Lee, and Chris Tan' in bioRxiv 2023, whereas [43] is the verifiable Bioinformatics 2024 article by Dotan et al. with the same title. This is a checkable attribution error and indicates that the bibliography has not been verified against authoritative sources. Please correct the citation or remove the duplicate, and verify all references for author, title, and DOI accuracy.","section":"References [5] and [43]"},{"comment":"The sentence 'Zhang et al. trained on 6,000 GPUs [30]' is implausible for the cited miTDS model [30], a fine-tuned BERT-based miRNA-mRNA interaction predictor with 110M parameters according to Table 1, and the claim is not supported by the cited article as far as the manuscript reports. Reporting an unsupported GPU count misrepresents the computational requirements of the reviewed methods and should be corrected with the actual training configuration or removed.","section":"§3.3"},{"comment":"The PRISMA checklist narrative states that 'risk of bias' and 'certainty of evidence' were assessed, but no risk-of-bias tool, no certainty assessment method, and no corresponding results appear in the methods or results sections. A scoping review under PRISMA-ScR does not require risk-of-bias assessment, but claiming it without reporting it is misleading. Please either implement and report these assessments or revise the PRISMA description to match what was actually done.","section":"Appendix A and §2.1"}],"minor_comments":[{"comment":"Please harmonize the database list between the abstract and the body; the current discrepancy (six databases in the abstract versus ten in the body) is confusing and undermines the reproducibility narrative.","section":"Abstract and §2.2"},{"comment":"The search dates '04/12/24' and '04/15/24' are ambiguous; use an unambiguous format such as '12 April 2024' to avoid confusion between April 12 and December 4.","section":"Appendix B"},{"comment":"The sentence beginning 'models have been developed to investigate the reusability and generalizability of cell-type annotation in single-cell RNA sequencing data' appears twice in the same paragraph; please remove the duplicate.","section":"§3.2, Cancer Research and Oncology"},{"comment":"The columns 'Model Derived Data' and 'Model Availability' are not labeled with two clear subheadings, making entries like 'Yes\\Yes' ambiguous; please separate data availability and model availability into distinct columns with explicit headers.","section":"Table 1"},{"comment":"Please verify the author lists of references [7] and [8]; the names 'Rohan Patel, Shalini Gupta, and Wei Wang' and 'Emily Brown, David Harris, and Ming Zhang' are generic and could not be confirmed against the cited sources during review.","section":"References [7] and [8]"},{"comment":"The PRISMA flowchart (Figure 1) is referenced, but the text does not report the initial number of records identified or the number remaining after deduplication; reporting these numbers would strengthen transparency.","section":"§2.4"}],"recommendation":"major_revision","confidential_remarks":"The paper has the right scope and the authors are well positioned to fix the issues, but the current version contains unverifiable search strings, an inconsistent database list, and at least one implausible computational-resource claim. I would ask the authors to supply the actual executed queries, per-database result counts, and verified reference metadata as supplementary material before reconsidering the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful overview of NLP-for-genomics work, but the reported search protocol is not reproducible as written, and there are concrete factual and referencing errors. The review is a good starting point for someone entering the field, not yet a trustworthy systematic map.\n\nWhat it does well: it organizes 26 recent studies into clear tables (models, goals, data, parameter sizes, tokenization, transformer type, availability), and the narrative around tokenization (k-mer, BPE, fixed nucleotide) and transformer architectures is a reasonable summary. The decision to highlight regulatory annotation prediction as a common downstream task is sensible and matches the literature. The PRISMA-style flowchart and checklist give the appearance of rigor, and the data availability statement is explicit.\n\nThe soft spots are real and load-bearing for a scoping review. Appendix B's queries cannot have been executed as printed: the ACM query contains bare '[Abstract: or]' and '[Abstract: or \"nlp\" or \"llm\" or]' fragments, and several strings end with dangling boolean operators. Section 2.2 also lists different databases in the abstract (PubMed, Medline, Scopus, Web of Science, Embase, ACM) versus the body (adding IEEE Xplore, Google Scholar, Semantic Scholar, Ovid MEDLINE, Ovid EMBASE). That ambiguity makes the record-identification step unverifiable, so the claim that these 26 studies are the relevant population is unsupported. Reference [5] duplicates the title of reference [43] but is attributed to generic authors ('John Smith, Anne Lee, Chris Tan')—a checkable bibliography error. And Section 3.3's claim that Zhang et al. trained on 6,000 GPUs for a miRNA-mRNA interaction model is implausible and likely a misreading of the source. None of these are fatal to the broad conclusion that transformer models are being applied to regulatory prediction, but they undermine the review's reliability as a definitive map.\n\nWho gets value? A newcomer wanting a quick orientation to methods and tasks in this space. A rigorous systematic reviewer should not rely on it until the search is corrected and verified. I'd send it to peer review with major revisions—the errors are fixable, and the field would benefit from a reliable version.","headline":"A useful but sloppy scoping review: the 26-study map is plausible, but the unreproducible search strategy and checkable errors mean it needs major revision before it can be treated as a reliable map.","tokens_in":16197,"tokens_out":2715,"would_cite":false,"duration_ms":24350,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A scoping review of 26 studies reports that tokenization and transformer models—the machinery of large language models—are now enhancing how genomic sequences are processed and understood, especially for predicting regulatory annotations.","keywords":["natural language processing","large language models","genomic sequencing","tokenization","transformer architecture","regulatory annotations","scoping review","DNA language models"],"falsifier":"A concrete way to test the central claim would be to build a benchmark that applies several of the surveyed tokenizers (k-mer, byte-pair encoding, fixed nucleotide) and transformer variants to the same set of regulatory-annotation tasks with identical evaluation protocols, and check whether transformer models consistently beat simpler sequence models; if a well-tuned CNN or k-mer counting baseline matches or exceeds the transformer results across tasks, the review's conclusion that transformers enhance genomic processing would not hold as stated.","tokens_in":15244,"feed_emoji":"🧬","tokens_out":5516,"duration_ms":46888,"temperature":0.7,"pith_summary":"This review seeks to map how natural language processing techniques, particularly large language models and transformer architectures, are being applied to genomic sequencing data, and whether recent literature supports the idea that these methods improve genomic analysis. Across 26 studies published between 2021 and April 2024, it finds that tokenization choices and transformer backbones are central to performance, with successful applications in predicting regulatory annotations such as transcription-factor binding sites, methylation sites, and chromatin accessibility. The review also documents that most datasets and models are publicly accessible and that computational demands vary widely, from heavy pretraining to lightweight feature extraction. If the synthesis is right, NLP-based genomic analysis is a practical and rapidly expanding route toward interpreting non-coding and regulatory parts of the genome, with downstream relevance for personalized medicine.","feed_headline":"Transformers decode regulatory DNA, review of 26 studies finds","feed_subtitle":"Tokenization plus attention models now predict binding sites, methylation, and chromatin state across the genome.","key_machinery":"The two load-bearing components the review tracks are tokenization and the transformer architecture. Tokenization converts raw DNA or RNA sequences into discrete units—most often k-mers (overlapping fixed-length substrings), byte-pair-encoding subwords, or fixed nucleotide fragments—so that sequence context can be treated like text. The transformer, typically a BERT-style encoder with multi-head self-attention, then captures long-range dependencies and contextual relationships among those tokens; a lightweight classifier is usually added on top to predict regulatory annotations such as binding sites, methylation status, chromatin accessibility, and RNA interactions. The review's argument is that this two-stage pipeline, borrowed from NLP, is what lets genomic models reach the reported predictive accuracy.","core_discovery":"The paper's central finding is that tokenization and transformer models enhance the processing and understanding of genomic data, particularly for predicting regulatory annotations. It surveys 26 studies selected from six databases and shows a consistent pattern: k-mer tokenization is the dominant preprocessing strategy, BERT-style transformers are the dominant architecture, and the most common downstream tasks are predictions of transcription-factor binding sites, methylation and CpG islands, enhancers, promoters, chromatin accessibility, and RNA interactions. The review reports performance metrics such as F1 scores above 0.8 for several models and notes that model and data availability are generally good, while computational requirements span a wide range.","pith_inferences":["If the reported metrics are taken at face value, it would be worth testing whether simple k-mer counts plus a conventional classifier already capture most of the signal attributed to transformers, since the review does not include negative controls or ablations.","The review's emphasis on DNA sequences and cancer-related tasks suggests that non-human genomes, rare diseases, and neurodegenerative conditions are under-explored niches where genomic language models could be tested next.","A direct extension would be a benchmark that runs several of the surveyed tokenizers and backbones on the same regulatory-annotation datasets to see whether the apparent advantage of transformers survives controlled comparison.","The trend toward multimodal data, such as genomic sequences combined with transcriptomic, proteomic, or imaging data, implies that the next generation of genomic language models may be evaluated on fusion tasks rather than sequence-only prediction."],"forward_implications":["K-mer tokenization is currently the default choice for genomic language models, so improvements in tokenization could shift performance across many downstream tasks.","Transformer-based models, especially BERT variants, are being used as feature extractors for regulatory annotation prediction, meaning pretrained genomic language models can be reused for multiple tasks.","Predictive targets cluster around regulatory biology—transcription-factor binding, methylation, chromatin accessibility, and non-coding RNA interactions—so the near-term payoff of NLP in genomics is chiefly in regulatory annotation rather than whole-genome interpretation.","Because most included models and datasets are publicly accessible, replication and fine-tuning on new genomic datasets should be feasible without starting from scratch.","Computational cost varies enormously: pretraining a genomic BERT can require days on many GPUs, while using existing embeddings as features is comparatively cheap, which affects who can enter the field."],"supporting_citations":[{"why":"Supplies the premise that tokenization choices affect transformer performance on biological sequences.","marker":"[5]"},{"why":"Provides the k-mer plus BERT baseline for regulatory element prediction that many included studies build on.","marker":"[12]"},{"why":"Demonstrates that task-specific pretraining improves DNA-protein binding prediction, anchoring the transformer-benefit claim.","marker":"[15]"},{"why":"Shows that pretrained transformer embeddings improve CpG island detection, a core regulatory annotation task.","marker":"[23]"},{"why":"Introduces a custom 6-mer tokenizer with taxonomic context and multiple transformer variants for methylation prediction with high reported F1.","marker":"[25]"},{"why":"Represents enhancer prediction by combining BERT embeddings with a 2D CNN for local feature extraction.","marker":"[13]"},{"why":"Uses fixed-nucleotide fragment tokenization for promoter prediction, supporting the review's account of tokenization variety.","marker":"[14]"},{"why":"Applies per-nucleotide tokenization with a transformer to detect oncoviral infections, supporting the review's claim of task diversity.","marker":"[16]"}],"fun_headline_variants":["Transformers unravel DNA regulatory code: 26-study review","NLP and LLMs decode genome: scoping review of 26 studies","k-mer tokenization plus BERT predicts binding sites and more","Genomic NLP review: attention models shine on regulatory DNA","26 studies show transformers read chromatin and binding sites"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's conclusion depends on the assumption that its database searches and eligibility criteria caught the relevant population of NLP-for-genomics studies and that the performance numbers reported by the 26 included papers are accurate and comparable across datasets; if major models were missed or metrics are inflated or inconsistent, the field would look more capable than it is.","fun_headline_variants_meta":{"raw":{"variants":["Transformers unravel DNA regulatory code: 26-study review","NLP and LLMs decode genome: scoping review of 26 studies","k-mer tokenization plus BERT predicts binding sites and more","Genomic NLP review: attention models shine on regulatory DNA","26 studies show transformers read chromatin and binding sites"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1337,"prompt_tokens":951,"completion_tokens":386,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":567,"completion_tokens_details":{"reasoning_tokens":302}},"tokens_in":567,"tokens_out":386,"duration_ms":3951,"temperature":1.0,"reasoning_tokens":302,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:32:13.370054+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete way to test the central claim would be to build a benchmark that applies several of the surveyed tokenizers (k-mer, byte-pair encoding, fixed nucleotide) and transformer variants to the same set of regulatory-annotation tasks with identical evaluation protocols, and check whether transformer models consistently beat simpler sequence models; if a well-tuned CNN or k-mer counting baseline matches or exceeds the transformer results across tasks, the review's conclusion that transformers enhance genomic processing would not hold as stated.","supporting_citations":[{"cited_title":"Effect of tokenization on transformers for biological sequences","cited_arxiv_id":null,"evidence_quote":"Supplies the premise that tokenization choices affect transformer performance on biological sequences."},{"cited_title":"Improving language model of human genome for dna–protein binding prediction based on task-specific pre-training","cited_arxiv_id":null,"evidence_quote":"Demonstrates that task-specific pretraining improves DNA-protein binding prediction, anchoring the transformer-benefit claim."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces a custom 6-mer tokenizer with taxonomic context and multiple transformer variants for methylation prediction with high reported F1."},{"cited_title":"A trans- former architecture based on bert and 2d convolutional neural network to identify dna enhancers from sequence information","cited_arxiv_id":null,"evidence_quote":"Represents enhancer prediction by combining BERT embeddings with a 2D CNN for local feature extraction."},{"cited_title":"Bert-promoter: An improved sequence-based predictor of dna promoter using bert pre-trained model and shap feature selection","cited_arxiv_id":null,"evidence_quote":"Uses fixed-nucleotide fragment tokenization for promoter prediction, supporting the review's account of tokenization variety."}],"review_version":1}