{"id":"5d8755ba-5155-4f92-8533-3ac6c558ec2c","arxiv_id":"2507.01494","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of deep learning crop pest classification that proposes a taxonomy and trend summary, but whose screening counts and reference attributions are internally inconsistent.","lead":"This paper surveys 37 studies on using deep learning to classify crop pests, organizing them by crop, pest type, model architecture, and dataset. It concludes that the field is shifting from CNNs toward hybrid and transformer models, while data imbalance and small pest detection remain open problems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PRISMA arithmetic in Section 2 cannot yield the claimed 37 studies: 355-21=334, 334-172=162 (not 166), and 162-121=41 (not 37); the systematic-coverage claim is therefore unreproducible until the inclusion set is reconstructed.","rationale":"The review's strongest claim is not a new experimental result but a map: 37 carefully selected studies and the taxonomy derived from them. The central condition for that claim is that the sample is exactly the product of a reproducible systematic search. Section 2 fails this condition on its face: the reported arithmetic gives 41 or 45 remaining studies, not 37. Because the review provides no query strings, no search dates, and no list of the 37 included studies, one cannot repair the count by checking an appendix; there is no appendix. The citation/table mismatches are directly visible evidence that the paper's internal indexing of its own sample is unreliable: Table 1 assigns the tomato YOLOv3 work to [17], which is a tea YOLOv8 paper, and the Insect-YOLO results are attributed to [6] rather than [39]. Undefined dataset initials (LLPD) add to the unverifiability. These are not merely cosmetic typos in a review whose conclusions are independently plausible; they attack the specific load-bearing assertion that the taxonomy is a systematic synthesis of a defined set. I agree with the reader's weak-assumption analysis, and the appropriate verdict is conditional: the narrative and many study summaries may be correct, but the selection layer must be corrected and made reproducible before the paper can be used as a reliable map. A corrected version that repairs the counts, lists the 37 studies, and fixes the reference attributions would satisfy the concern. I found no evidence of bad faith; the failure is one of verifiability, not integrity.","tokens_in":19976,"tokens_out":5521,"duration_ms":64850,"concrete_test":"Reconstruct the selection chain and the 37-item inclusion list. (1) Re-derive the PRISMA counts from the stated search and inclusion criteria; if 355-21=334, 334-172=162 (not 166), and 162-121=41 (not 37) cannot be reconciled with the announced final count, the screening claim fails. (2) Enumerate the primary studies actually cited in Sections 4-8 and Tables 1-13 and compare with the announced 37; the number and identity of included studies should be exactly recoverable. (3) Verify the attributions for at least [17] and [6] against the cited PDFs; if these entries mismatch, check all table rows for systematic misattribution. The minimum settlement is a PRISMA-style supplement listing all 37 studies with unique IDs and exclusion reasons.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2's screening arithmetic cannot produce the announced final set. The paper reports 355 records, 21 duplicates, 172 title/abstract exclusions, 166 records left for further review, and 121 full-text exclusions. Independently: 355-21=334; 334-172=162, not 166; and 162-121=41, not 37 (the paper's own stated 166-121=45). No consistent reading of the stated numbers yields 37 unless an exclusion count is changed. This is load-bearing because the abstract and conclusion stake the entire value of the review on 37 carefully selected studies published between 2018 and 2025, and on the taxonomy in Sections 3-8 being representative of that set. If the sample cannot be reconstructed from the reported search, the systematic-coverage claim is unverifiable. The trend narrative (CNNs to hybrid/transformer models; imbalance, small-pest detection, generalization, edge deployment) may survive correction, but it would then be a selective narrative, not the systematic map the paper claims. The problem is compounded by citation/table mismatches that make even the reviewed set hard to identify: Table 1 attributes a tomato YOLOv3 study to [17], which is the tea YOLOv8 paper; Insect-YOLO results are attached to [6] while the actual Insect-YOLO paper is [39]; and Section 7 introduces LLPD without defining it. These are all verifiable in the text and point to a single underlying issue: the selection process and its products are not reproducible enough to carry the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a literature review of deep learning methods for crop pest classification, claiming to survey 37 systematically selected studies published between 2018 and 2025. It organizes the literature by crop, pest type, model architecture (CNN, vision transformer, hybrid, object detection), datasets, and technical challenges, and argues that the field has shifted from CNN-only models toward hybrid and transformer-based models. The review concludes by discussing challenges such as class imbalance, small-pest detection, generalization, and edge deployment, and suggests future research directions.","tokens_in":20128,"tokens_out":5688,"duration_ms":57865,"significance":"As a review, the paper's value depends entirely on the reliability and representativeness of its literature selection and on the correctness of its attributions. If the systematic selection were as described, the taxonomy and tables would provide a useful map of the field. The paper has strengths: it covers a wide range of studies, identifies key benchmark datasets (IP102, PlantVillage), offers a structured taxonomy, and outlines important open challenges, including emerging techniques such as quantum-inspired CNNs and diffusion-based augmentation. However, the internal inconsistencies noted below (PRISMA arithmetic and citation/table mismatches) undermine confidence in the review's central claim of systematic coverage, as the set of 37 studies cannot be reconstructed.","major_comments":[{"comment":"The reported screening numbers are internally inconsistent. 355-21=334; 334-172=162, not 166; and 162-121=41 (or 166-121=45), not 37. No recalculation using the stated figures yields the final set. Furthermore, the manuscript never identifies which 37 studies are included. These two problems together make the systematic-coverage claim unreproducible, and since the abstract and conclusion explicitly stake the review's value on '37 carefully selected studies,' this is a load-bearing issue.","section":"Section 2"},{"comment":"The tomato YOLOv3 study is attributed to reference [17], but Section 4.2 attributes it to Liu and Wang [14], and reference [17] is the tea YOLOv8 paper. Similarly, Section 4.1 describes reference [12] as a CNN-based rice pest classifier, whereas the cited reference is a transformer-based pest detector (used as such in Sections 6.2 and 8.2). These misattributions mean the tables and text do not reliably allow readers to identify the reviewed studies.","section":"Table 1"},{"comment":"The sentence 'In addition to established benchmark datasets like IP102, PlantVillage, and LLPD' introduces 'LLPD' without ever defining it. Since Section 7 is meant to describe benchmark datasets, this omitted definition is a substantive gap. Additionally, the subsections are titled 'Xei-1' and 'Xei-2,' which should be 'Xie-1' and 'Xie-2' after the author's name.","section":"Section 7.5"}],"minor_comments":[{"comment":"'baseline CNN accuracy was 49' should read '49%' for clarity.","section":"Section 7.1"},{"comment":"The term 'HPMA-ViT' is used without definition; the later Section 6.3 refers to 'HP-MHA' (hybrid pooled multi-head attention), so the acronym should be introduced consistently.","section":"Section 3"},{"comment":"The statement that 'the IP102 benchmark... has exposed the difficulty of fine-grained classification with a baseline accuracy below 50% (ResNet-50)' conflicts with Section 7.1's 'baseline CNN accuracy was 49' only in presentation; the numbers should be presented with a consistent notation.","section":"Section 4.1"},{"comment":"Several table captions list 'Y ear' for 'Year'; this typo should be corrected.","section":"Tables 3-6"}],"recommendation":"major_revision","confidential_remarks":"This manuscript has the potential to be a useful review, but the systematic methodology needs to be repaired and verified before publication. The PRISMA arithmetic error is basic and suggests that the screening numbers were not checked; the editors may want a methodological reviewer to verify the revised counts and the list of included studies. I would not consider acceptance before these issues are resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. This is a review, so novelty is zero by definition, and the paper doesn't pretend otherwise. What it does well: it gives a newcomer a genuinely organized map of crops (rice, tomato, peanut, tea, cactus), pest types, model families (CNN, ViT, hybrid, YOLO variants), and datasets (IP102, Xie datasets, PlantVillage), plus a fair summary of known bottlenecks: class imbalance, small pest detection, field robustness, edge deployment. The trend narrative—CNNs giving way to hybrids and transformers—is plausible and matches the broader survey literature. The tables, typos aside, are a useful quick reference.\n\nThe soft spots are real and load-bearing. Section 2's PRISMA arithmetic doesn't reconstruct the 37 studies: 355 minus 21 duplicates is 334; 334 minus 172 title/abstract exclusions is 162, not 166; and 162 minus 121 full-text exclusions is 41, not 37. The paper's own 166 minus 121 also gives 45, not 37. No set of stated numbers yields 37. That means the systematic-coverage claim—the entire raison d'être of a review like this—is unverifiable. The authors need to redo the screening counts or explain the discrepancy. Compounding this, Table 1 attributes a tomato YOLOv3 study to reference [17], which is actually a tea pest YOLOv8 paper; Section 7 name-drops LLPD without defining it; and the search strategy lacks dates and full query strings. These are fixable, but they're exactly the errors that make a review untrustworthy as a map.\n\nThe central trend conclusions probably survive correction—the direction of the field is not in doubt. But as it stands, the paper is a selectively organized review wearing a systematic review's clothes. A careful referee would ask for minor-to-moderate revisions: fix the screening numbers, reconcile the table citations, define every dataset acronym, specify the search protocol. None of that requires new experiments, so it's within reach.\n\nWho is this for? A graduate student or practitioner wanting a first bird's-eye view of deep learning for crop pest classification would get value, with the caveat to check primary sources. I'd bring it to a reading group as a case study in systematic-review reporting. I don't think I'd cite it in my own work until the numbers are fixed. That said, it deserves a serious referee: the topic is timely, the coverage is substantial, and the flaws are correctable.","headline":"Useful newcomer's map of crop pest deep learning, but the systematic selection claim is not reproducible from the paper's own PRISMA arithmetic.","tokens_in":20772,"tokens_out":2235,"would_cite":false,"duration_ms":24706,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 37-study review maps crop pest classification from plain CNNs to hybrid and transformer models.","keywords":["crop pest classification","deep learning","vision transformers","convolutional neural networks","hybrid CNN-transformer","YOLO object detection","IP102 dataset","agricultural AI"],"falsifier":"Recompute the Section 2 screening numbers: 355 records minus 21 duplicates is 334, 334 minus 172 title-or-abstract exclusions is 162 rather than the stated 166, and 162 minus 121 full-text exclusions is 41 rather than 37. Counting the studies actually cited and comparing that number to the claimed 37 would settle whether the systematic-coverage claim holds; a direct CNN-versus-transformer-versus-hybrid benchmark on IP102 under equal compute would test the trajectory claim.","tokens_in":19636,"feed_emoji":"🐛","tokens_out":8506,"duration_ms":87159,"temperature":0.7,"pith_summary":"This review tries to establish a map of how deep learning has been used to classify insect pests on crops between 2018 and 2025, based on 37 selected studies. Its central claim is that the field follows a clear trajectory: early work relied on convolutional neural networks, while newer work shifts to vision transformers and hybrid CNN-transformer models that report higher accuracy and better contextual understanding. A sympathetic reader would care because the review organizes the literature into a usable taxonomy by crop, pest type, model architecture, dataset, and challenge, and in doing so names the benchmarks and open problems that determine whether automated pest monitoring can move from lab to field. If the review is right, the practical direction is set: YOLO-family detectors for locating pests, hybrid architectures for fine-grained classification, and larger, more balanced datasets as the main bottleneck.","feed_headline":"Deep learning pest ID shifts from CNNs to hybrids","feed_subtitle":"A 37-study review maps crops, pests, architectures, and datasets for automated insect classification.","key_machinery":"The machinery is the review's five-axis taxonomy (crop, pest type, AI technique, dataset, challenge) together with a systematic selection protocol that filters 355 initial records down to 37 studies. The taxonomy is the load-bearing object: it converts individual papers into comparable entries and is what lets the review claim that a CNN-to-hybrid/transformer trajectory and a fixed set of challenges are genuine trends rather than isolated results.","core_discovery":"On its own terms, the paper establishes that published pest-classification research from 2018 to 2025 is not a random collection of case studies but a structured field with five recurring dimensions: crop, pest type, technique, dataset, and challenge. It reports that CNN classifiers dominated early work and often exceed 90% accuracy on curated datasets, while vision transformers and hybrid CNN-transformer models now achieve comparable or higher accuracy, particularly where global context or fine-grained differences matter, citing examples such as EViTA on peanut pests and CactiViT on cactus cochineal. At the same time, it documents persistent bottlenecks: the IP102 benchmark's imbalanced 102-class distribution keeps baseline accuracy below 50%, small pests occupy too few pixels for standard detectors, field conditions degrade accuracy, and models do not transfer well across crops or regions. The review presents these findings through a taxonomy and summary tables rather than new experiments.","pith_inferences":["The review does not itself run a controlled comparison; a direct test would be to compare pure-CNN, pure-transformer, and hybrid models on the same public benchmark (for example IP102) under identical training budgets, since the trajectory claim predicts hybrids should win.","The small-pest challenge points to a concrete research program the review describes but does not develop: pairing super-resolution or patch-based high-resolution inputs with attention modules, tested on sticky-trap or drone imagery.","The reported screening arithmetic is internally inconsistent (334 minus 172 leaves 162, and subtracting 121 leaves 41, not 37), so the systematic-coverage claim should be checked against the actual reference list before the 37-study count is relied upon.","If the trajectory is correct, a practical consequence is that agricultural monitoring systems will standardize on a two-stage recipe: a YOLO-family detector for localization and counting plus a hybrid classifier for species-level identification, deployed on edge hardware."],"forward_implications":["Researchers entering the area get a checklist: choose a crop, pest type, architecture family, dataset, and challenge to address, with the review's tables mapping who has already done what.","If the trajectory claim holds, new high-accuracy systems will be hybrids or transformer-enhanced detectors, and pure CNN classifiers on small datasets will serve as baselines rather than endpoints.","Benchmarking on IP102 remains a stricter test than custom crop-specific datasets, and the sub-50% baseline accuracy marks the fine-grained generalization gap.","Deployment claims are becoming central: smartphone and edge-device models such as CactiViT and lightweight YOLO variants are treated as realistic, so efficiency is now part of the correctness story.","Data scarcity, not architecture, is the limiting factor, with self-supervised pre-training and augmentation as the emerging mitigations."],"supporting_citations":[{"why":"Supplies IP102, the large imbalanced 102-class pest benchmark that anchors the dataset and long-tail discussion.","marker":"[8]"},{"why":"Introduces the Vision Transformer architecture that the review identifies as the field's newer direction.","marker":"[5]"},{"why":"Provides the prior systematic review of deep-learning insect detection that this review extends and compares against.","marker":"[9]"},{"why":"Exemplifies the hybrid CNN-plus-transformer architecture and reports accuracy gains over standalone models on peanut pests.","marker":"[16]"},{"why":"Demonstrates a transformer running on smartphones for a niche cactus pest, supporting the small-data and deployment claims.","marker":"[7]"},{"why":"Exemplifies YOLO-family detectors engineered for small pests in cluttered field images.","marker":"[13]"},{"why":"Supports the hybrid trajectory with state-of-the-art results on a large-scale multi-class pest recognition task.","marker":"[11]"},{"why":"Shows a lightweight CNN with attention can cut model size, anchoring the efficiency and edge-deployment discussion.","marker":"[23]"},{"why":"Shows self-supervised pre-training can reduce labeling needs, relevant to the data-scarcity challenge.","marker":"[24]"}],"fun_headline_variants":["Pest ID review: CNNs yield to hybrid models","37-study review maps crop pest AI's shift to transformers","Vision transformers gain ground in pest classification","Pest classification review: hybrids edge out CNNs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 37 studies described are the representative output of the systematic search whose arithmetic is reported in Section 2; if that count or the selection is wrong, the review's claim to map the field loses its grounding.","fun_headline_variants_meta":{"raw":{"variants":["Pest ID review: CNNs yield to hybrid models","37-study review maps crop pest AI's shift to transformers","Vision transformers gain ground in pest classification","Pest classification review: hybrids edge out CNNs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1273,"prompt_tokens":899,"completion_tokens":374,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":312}},"tokens_in":515,"tokens_out":374,"duration_ms":5308,"temperature":1.0,"reasoning_tokens":312,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:50:19.657151+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the Section 2 screening numbers: 355 records minus 21 duplicates is 334, 334 minus 172 title-or-abstract exclusions is 162 rather than the stated 166, and 162 minus 121 full-text exclusions is 41 rather than 37. Counting the studies actually cited and comparing that number to the claimed 37 would settle whether the systematic-coverage claim holds; a direct CNN-versus-transformer-versus-hybrid benchmark on IP102 under equal compute would test the trajectory claim.","supporting_citations":[{"cited_title":"IP102: A Large-Scale Benchmark Dataset for Insect Pest Recognition","cited_arxiv_id":null,"evidence_quote":"Supplies IP102, the large imbalanced 102-class pest benchmark that anchors the dataset and long-tail discussion."},{"cited_title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","cited_arxiv_id":null,"evidence_quote":"Introduces the Vision Transformer architecture that the review identifies as the field's newer direction."},{"cited_title":"A systematic review on automatic insect detection using deep learning","cited_arxiv_id":null,"evidence_quote":"Provides the prior systematic review of deep-learning insect detection that this review extends and compares against."},{"cited_title":"Pest Detection and Classification in Peanut Crops Using CNN and EViTA Algorithms","cited_arxiv_id":null,"evidence_quote":"Exemplifies the hybrid CNN-plus-transformer architecture and reports accuracy gains over standalone models on peanut pests."},{"cited_title":"CactiViT: Image-based smartphone application and transformer network for diagnosis of cactus cochineal","cited_arxiv_id":null,"evidence_quote":"Demonstrates a transformer running on smartphones for a niche cactus pest, supporting the small-data and deployment claims."},{"cited_title":"YOLOv7-PSAFP: Crop pest and disease detection based on improved YOLOv7","cited_arxiv_id":null,"evidence_quote":"Exemplifies YOLO-family detectors engineered for small pests in cluttered field images."},{"cited_title":"Pest-ConFormer: A Hybrid CNN-Transformer Architecture for Large-Scale Multi-Class Crop Pest Recognition","cited_arxiv_id":null,"evidence_quote":"Supports the hybrid trajectory with state-of-the-art results on a large-scale multi-class pest recognition task."},{"cited_title":"GA-GhostNet: A Lightweight CNN Model for Identifying Pests and Diseases Using a Gated Multi-Scale Coordinate Attention Mechanism","cited_arxiv_id":null,"evidence_quote":"Shows a lightweight CNN with attention can cut model size, anchoring the efficiency and edge-deployment discussion."},{"cited_title":"Self-supervised learning improves classification of agriculturally important insect pests in plants","cited_arxiv_id":null,"evidence_quote":"Shows self-supervised pre-training can reduce labeling needs, relevant to the data-scarcity challenge."}],"review_version":1}