{"id":"fa4aedf3-aad6-4be9-b83e-f0695f6cdf98","arxiv_id":"2507.18532","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"COT-AD provides over 25,000 cotton crop images with about 5,000 annotated for detection, segmentation, and disease classification across the full growth cycle.","lead":"Researchers compiled COT-AD, a dataset of more than 25,000 images of cotton crops from drone and DSLR cameras, with about 5,000 images labeled for disease classification, detection, and segmentation. The dataset aims to give computer vision models a larger, multi-task resource for monitoring cotton health and supporting precision agriculture.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2 overclaims: it marks DSLR Disease as supporting detection, but Sec. 16 provides only folder-level disease labels; DSLR Segment lists 12 classes vs. 9 elsewhere; synthesis counts (14k+11k) conflict with the 25k total. The 'comprehensive' claim is not internally supported.","rationale":"The reader's CONDITIONAL verdict is appropriate. My stress-test sharpens the condition: the paper must reconcile its dataset split table with its dataset organization section and actual release. The strongest evidence in favor of the dataset is that it appears to exist (multiple distribution channels, ~310 GB, DOI), and basic experiments (classification 83.37%, StyleGAN2-ADA synthesis metrics, restoration tables) are reported. However, the load-bearing claim is the capability matrix (Tables 1-2), and the paper's own Section 16 contradicts it for detection and for the 12-class DSLR segment set. These are factual inconsistencies, not mere underreporting: if DSLR Disease images have no detection labels, then 'Detection ✓' for that split is unsupported, and the dataset supports detection only on aerial images (1 class). The 'weed analysis' annotation claim in the abstract is likewise unsupported. These issues can be fixed by updating Table 2, adding label inventories, and correcting the image/class counts; hence rejection is not warranted, but acceptance requires those corrections, matching the CONDITIONAL verdict. I disagree with the reader's locating the weakest assumption purely in annotation accuracy; the more basic risk is that claimed label types and counts do not exist as described. Hence 'partial' agreement.","tokens_in":12108,"tokens_out":5724,"duration_ms":54446,"concrete_test":"Download the dataset from Kaggle or IEEE DataPort and run an inventory script that (a) counts images and label files per split; (b) checks whether the 'DSLR Disease' folder contains any detection annotations (YOLO .txt or JSON bounding boxes) as Table 2 claims; (c) verifies the 'DSLR Segment' split exists with exactly 100 images and 12 class labels; (d) sums unique images across Drone, DSLR Disease, DSLR Segment, and synthesis splits and compares to the claimed 'over 25,000' and '5,000 annotated'; and (e) looks for any weed label files. If DSLR Disease has no target-level labels, or segment/synthesis counts don't reconcile, the 'comprehensive' task-support claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that COT-AD is a comprehensive dataset whose annotations support classification, detection, segmentation, enhancement, and synthesis. The paper's own dataset documentation does not substantiate that task matrix. (1) Table 2 assigns a detection checkmark to the DSLR Disease split (3,231 images), but Section 16 describes the DSLR data only as folder-organized class labels for disease classification, with no bounding-box or YOLO label files for DSLR images. Detection is only described for aerial images (single-class cotton crop). (2) Table 2 lists a DSLR Segment split of 100 images with 12 classes, but Section 4.3 defines 9 disease classes and Table 4 enumerates 11 categories; no other mention of a 12-class DSLR segmentation set exists. (3) The abstract says 'over 25,000 images ... with 5,000 annotated images,' while Table 2 gives a synthesis subset of 14,000 drone + 11,000 DSLR images for enhancement/synthesis; if those 25,000 are disjoint from the 5,131 annotated images, the total exceeds 25,000, and if overlapping, the count of 'annotated' images is ambiguous. (4) The abstract claims annotations cover 'weed analysis,' but the annotation procedure (Sec. 3) only labels cotton crops, and no weed labels are described anywhere. The reader's concern about annotation accuracy is real but secondary: before label quality can be evaluated, the labels must exist and be consistently counted for every claimed task.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces COT-AD, a cotton analysis dataset combining drone-based aerial imagery and handheld DSLR close-up images, and reports baseline experiments for disease classification, detection, segmentation, image enhancement, restoration, and synthesis. The central contribution is the dataset itself: public release, large volume, and a claimed multi-task annotation suite. The baseline results demonstrate that off-the-shelf models can be trained and evaluated on parts of the data. However, the manuscript contains several internally inconsistent dataset statistics and documentation gaps, and the supplementary organization does not support every checkmark in the claimed task matrix. These issues affect the paper's central claim of presenting a comprehensive, multi-task dataset and need to be resolved before the contribution can be accepted as stated.","tokens_in":12405,"tokens_out":4162,"duration_ms":46311,"significance":"If its documentation were made internally consistent, COT-AD would be a useful resource for the agricultural computer-vision community: it is publicly released with DOIs and a project page, covers two farms across a full growth cycle, uses multiple capture altitudes, and provides YOLO-format labels for aerial detection and segmentation. The reported baselines, including VGG19 disease classification at 83.37% test accuracy and StyleGAN2-ADA synthesis metrics, indicate that the data are usable with existing methods. The main weakness is not the experimental execution but the gap between the advertised task coverage and the documented annotation artifacts. In its current form, the paper overclaims detection support for the DSLR disease split and weed-label coverage, and the image-count arithmetic is not reproducible from Table 2. These are fixable documentation issues, but they are load-bearing for the dataset's stated contribution.","major_comments":[{"comment":"Table 2 assigns a detection checkmark to the DSLR Disease split (3,231 images), but §16 documents only folder-level disease labels for DSLR images and describes detection labels exclusively under the 'Aerial Images for Detection and Segmentation' directory. No YOLO .txt or bounding-box labels are described for DSLR images. The detection-support claim for the DSLR split is therefore not supported by the dataset documentation and should either be corrected in the task matrix or implemented and documented.","section":"Table 2; Supplementary §16"},{"comment":"The DSLR Segment row of Table 2 lists 100 images and 12 classes, but §4.3 and §16 define 9 disease classes, while Table 4 enumerates 11 categories (four leaf, three boll, two flower, two bug). No 12-class DSLR segmentation set is described anywhere in the supplementary organization. The class count, the image count, and the existence and format of DSLR segmentation masks must be clarified and made consistent across Tables 1, 2, and 4.","section":"Table 2; §4.3; Table 4; Supplementary §16"},{"comment":"The image-count arithmetic does not close. Table 2 sums to 5,131 annotated images (1,800 + 3,231 + 100), while the abstract states 'over 25,000 images ... with 5,000 annotated images.' The Synthesis split alone accounts for 25,000 images (14,000 + 11,000), which implies either a total well above 25,000 or overlap with the annotated splits. A single consistent accounting of total images, annotated images, and synthesis-subset images is needed.","section":"Abstract; Table 2; Supplementary §8"},{"comment":"The abstract and §9 claim that annotations cover 'weed analysis' and insights into 'weed competition,' but the annotation procedure in §3 labels only cotton crops, and §16 describes no weed label files. No weed annotations are documented anywhere. The authors should either release weed labels and describe them in §16, or remove weed-analysis coverage from the abstract and §9, since it is part of the claimed task coverage.","section":"Abstract; §3; §9; Supplementary §16"},{"comment":"The annotation-quality evidence is missing. Section 3 states that 'Each image was meticulously annotated,' but the paper reports no inter-annotator agreement, no label-error estimate, and no per-class image counts for the disease classification split. Given that the utility of the dataset, and the 83.37% VGG19 classification result, depends on label quality, the authors should provide basic label statistics and a quality-assessment measure.","section":"§3; §4.3"}],"minor_comments":[{"comment":"The text says 'we are using only 100 image points to train the classifier,' which is ambiguous given a 3,231-image disease split; please state the actual training-set size per class and reconcile it with the 70:15:15 split described in the same section.","section":"§4.3; Table 6"},{"comment":"The sentence 'we have used CLIP-RC [17] used YOLOv11 [5]' is grammatically incomplete, and reference [5] is a YOLO-based weed benchmark rather than the YOLOv11 model; the sentence and citation need correction.","section":"§4.2; Fig. 5"},{"comment":"The class-count column is inconsistent across tables: Table 1 lists 9 classes for COT-AD, Table 2 lists 12 classes for the DSLR Segment split, and Table 4 lists 11 categories; please unify these numbers or explain how they differ.","section":"Table 1; Table 2; Table 4"},{"comment":"The dataset volume is given as approximately 310 GB in §8 and approximately 308 GB in §15; the two figures should be reconciled.","section":"Supplementary §8; §15"}],"recommendation":"major_revision","confidential_remarks":"The dataset itself appears to be a real and potentially valuable contribution, and the inconsistencies identified here are correctable through revised documentation rather than new experiments. The main risk is that the authors may not actually possess detection labels for the DSLR disease images or weed annotations; if that is the case, the corresponding claims must be retracted even if the dataset remains useful for classification, aerial detection, and segmentation. The paper would benefit from a single consolidated dataset-statistics table."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know that the stress-test note holds up. The paper's central \"comprehensive\" claim is not supported by its own documentation, and the inconsistencies are concrete, not style complaints. The dataset is a genuine new resource: first to pair aerial and DSLR imagery for cotton, two farms, full season, 25k+ images, 310GB, with baseline experiments showing it can drive classification, detection, segmentation, synthesis, and restoration. That is real collection effort and useful baseline evidence. But the paper's own tables contradict each other. Table 2 marks DSLR Disease as supporting detection, yet Section 16 describes only folder-level disease labels for DSLR images; detection labels exist only for aerial images. DSLR Segment is listed as 12 classes, while Table 4 enumerates 11 categories and Section 4.3 defines 9. The synthesis split (14k drone + 11k DSLR) plus the annotated splits (1,800 + 3,231 + 100) cannot both be subsets and add up to 25k. The abstract promises weed analysis, but no weed labels appear anywhere. These are not cosmetic; they make it impossible to know what you are actually downloading. The reader's concern about annotation quality is secondary but real: no IAA, no per-class counts, no error rates. The manual annotation claim is asserted without evidence. Baselines are thin but adequate for a dataset paper; the real issue is that the dataset's task matrix is unverified. The paper is short and workshop-flavored, but the data collection and initial experiments show serious effort. This deserves a serious referee, not a desk reject, because the dataset could be valuable if the documentation is fixed. The referee should require the authors to reconcile the counts, define the actual annotation format per split, and either provide weed labels or drop that claim. As written, my verdict would be conditional-to-reject, but the underlying resource is worth the referee's time.","headline":"A large, potentially useful cotton dataset undermined by internally inconsistent claims about what it contains.","tokens_in":13009,"tokens_out":2200,"would_cite":false,"duration_ms":24295,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces COT-AD, over 25,000 drone and DSLR images of cotton with about 5,000 annotated, and calls it the first dataset to support classification, detection, segmentation, restoration, enhancement, and synthesis for the crop.","keywords":["cotton crop dataset","precision agriculture","disease classification","object detection","semantic segmentation","image restoration","generative synthesis","aerial imagery"],"falsifier":"Take a random sample of the 5,000 annotated images, have two independent annotators relabel them, and measure agreement: if agreement on the nine disease and pest classes is low, the annotation-quality claim collapses. A second check: train the same VGG19 pipeline on COT-AD and evaluate it on independently photographed cotton images from a different region or a different season; a drop toward chance on classes such as boll rot or leaf reddening would show the two-farm, one-season collection does not generalize as claimed.","tokens_in":11925,"feed_emoji":"🌱","tokens_out":10978,"duration_ms":104827,"temperature":0.7,"pith_summary":"Intended as a one-stop resource for computer-vision work on cotton, this paper introduces COT-AD: more than 25,000 drone and DSLR images of cotton fields and plants, captured across a six-month growing season on two farms, with about 5,000 of those images annotated for detection, segmentation, and disease class. The authors' central claim is that this combination fills a real gap, because every earlier cotton dataset in their comparison is smaller, single-modality, and built for one task, whereas COT-AD is designed to support six at once: classification, detection, segmentation, restoration, enhancement, and generative synthesis. A reader should care because precision agriculture depends on data that reflects the full crop cycle and both the field-level and the plant-level view, and no existing cotton benchmark, if the comparison is right, offered that in one collection.","feed_headline":"25,000 cotton images cover the whole crop cycle in one dataset","feed_subtitle":"Aerial and close-up photos with 5,000 annotations combine classification, detection, and segmentation for farm AI.","key_machinery":"The central object is the dataset's structure itself, which is what carries the argument. The aerial half is partitioned into four time parts (first two months, third month, fourth month, fifth-sixth months), so crop age is encoded by folder structure, and each part stores images, YOLO-format detection labels, binary segmentation masks, and YOLO-format segmentation labels side by side; the DSLR half is organized into per-class folders under Leaf, Cotton Boll, and Bugs, covering nine disease and pest classes. This month-by-month design is the mechanism that lets one collection serve both spatial tasks (detect and segment crops in the field) and temporal tasks (track disease onset and spread across the season), and the paper supplies month-wise interpretation tables that tie each stage to the appearances visible in the data. Six demonstration pipelines then carry the validation claim: VGG19 for disease classification, CLIP and BioCLIP linear probing, Deep Spectral Method and CLIP-RC for segmentation, StyleGAN2-ADA for synthesis, and RUSIR and BRGM for restoration.","core_discovery":"Stated the way the authors would state it: cotton lacks the kind of rich, multi-task dataset that other crops enjoy, and COT-AD exists to close that gap. The dataset bundles a drone-captured aerial set with images from a Canon EOS 80D DSLR, both taken on the same farms across the same season, so that field-scale structure and close-up disease symptoms come from one continuous observation campaign. Aerial frames carry single-class YOLO-format detection labels and binary segmentation masks, organized into month-based parts that encode crop age; the DSLR set groups 3,231 images into nine classes spanning leaf diseases, boll diseases, and two insect pests. To show the data are usable, the paper runs six task families on it and reports a VGG19 disease-classification accuracy of 83.37%, CLIP and BioCLIP zero-shot and linear-probing scores, unsupervised and CLIP-guided segmentation, StyleGAN2-ADA synthesis, and two generative restoration pipelines.","pith_inferences":["Because crop age is encoded in the folder structure, a shortcut probe is natural: a classifier trained with month labels blurred (or images shuffled across parts) would reveal how much of the 83.37% accuracy relies on temporal stage cues rather than on disease appearance.","The paper never reports per-class instance counts; a distributional audit would tell users which of the nine classes are trainable and which are too sparse, and the dataset's usability for rare pests like the Red Cotton Bug hangs on this.","Two farms, one region, one season: fine-tuning on COT-AD and testing on an independent cotton image set from another region would settle whether the benchmark transfers or captures site-specific lighting, soil, and variety.","A forward path the paper does not take: since drone and DSLR views cover the same plants, one could train a cross-modality model that enhances low-altitude frames to DSLR-level detail, fusing the two halves the dataset was built to join."],"forward_implications":["Any cotton-disease classifier can now be tested against a common benchmark, with 83.37% VGG19 accuracy as the first number to beat.","Because the aerial images are organized by crop month, a model can learn to infer growth stage from field appearance, which enables scheduling of scouting or spraying before close-up symptoms are visible.","The same acquisition campaign provides paired drone and DSLR views, so restoration and enhancement methods can be trained and evaluated on realistic field imagery rather than on generic photos.","Synthetic images from StyleGAN2-ADA trained on the dataset can augment underrepresented disease classes during training, helping classifiers in data-scarce settings.","Feeding the dataset into the demonstrated YOLO-based detection pipeline gives a ready-made field-monitoring workflow for weed competition and plant counting, a direct corollary of the annotations provided."],"supporting_citations":[{"why":"Supplies the closest prior aerial work: low-altitude UAV cotton imagery used for yield estimation and segmentation, which COT-AD extends toward detection and classification.","marker":"[1]"},{"why":"Provides the augmented Cotton Leaf Disease Detection dataset used in the comparison table and as an external test set for classification and enhancement.","marker":"[2]"},{"why":"A cotton leaf image dataset that anchors the comparison table's claim that existing sets are small, single-task, and leaf-only.","marker":"[3]"},{"why":"The cotton leaf disease dataset used as another scope baseline and in cross-dataset classification and enhancement runs.","marker":"[4]"},{"why":"YOLOWeeds, the cotton-production weed detection benchmark that supplies the YOLO detector framework used for the detection experiments.","marker":"[5]"},{"why":"Deep Spectral Method, the unsupervised segmentation baseline applied to the aerial images to demonstrate the segmentation annotations.","marker":"[14]"},{"why":"CLIP, the vision-language backbone for the zero-shot and linear-probing classification results that benchmark the dataset's learnability.","marker":"[18]"},{"why":"The VGG19 architecture comparison that the disease classifier is built on, producing the 83.37% test accuracy.","marker":"[20]"},{"why":"Bayesian image reconstruction using generative models (BRGM), one of the two restoration pipelines applied to drone and DSLR splits.","marker":"[21]"},{"why":"Robust unsupervised StyleGAN image restoration (RUSIR), the framework used for denoising, deartifacting, upsampling, and inpainting evaluations.","marker":"[22]"}],"fun_headline_variants":["Cotton gets its own vision dataset: 25k images and 5k labels","Aerial to close-up: one cotton dataset for detection and segmentation","First cotton dataset linking field-scale and disease-level imagery","25k cotton images: drone views plus DSLR close-ups for six tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the unverified quality of the manual labels: the paper reports that every image was 'meticulously annotated' (Section 3) yet gives no inter-annotator agreement, no error rate, and no per-class counts, and all imagery comes from just two farms in one growing season (Sections 3 and 11); noisy or unrepresentative labels would invalidate both the dataset's utility and the 83.37% accuracy baseline.","fun_headline_variants_meta":{"raw":{"variants":["Cotton gets its own vision dataset: 25k images and 5k labels","Aerial to close-up: one cotton dataset for detection and segmentation","First cotton dataset linking field-scale and disease-level imagery","25k cotton images: drone views plus DSLR close-ups for six tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000794,"raw_usage":{"total_tokens":3446,"prompt_tokens":841,"completion_tokens":2605,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":2527}},"tokens_in":457,"tokens_out":2605,"duration_ms":17921,"temperature":1.0,"reasoning_tokens":2527,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:31:09.121259+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the 5,000 annotated images, have two independent annotators relabel them, and measure agreement: if agreement on the nine disease and pest classes is low, the annotation-quality claim collapses. A second check: train the same VGG19 pipeline on COT-AD and evaluate it on independently photographed cotton images from a different region or a different season; a drop toward chance on classes such as boll rot or leaf reddening would show the two-farm, one-season collection does not generalize as claimed.","supporting_citations":[{"cited_title":"COT-AD: Cotton Analysis Dataset","cited_arxiv_id":"2507.18532","evidence_quote":"Supplies the closest prior aerial work: low-altitude UAV cotton imagery used for yield estimation and segmentation, which COT-AD extends toward detection and classification."},{"cited_title":"While the first two focus on leaf images, the third provides a broader scope by including both leaf and whole-plant images","cited_arxiv_id":null,"evidence_quote":"Provides the augmented Cotton Leaf Disease Detection dataset used in the comparison table and as an external test set for classification and enhancement."},{"cited_title":"It includes high-resolution aerial imagery collected at altitudes of [10m, 15m, and 115m] and detailed close-up images from handheld DSLR cameras","cited_arxiv_id":null,"evidence_quote":"A cotton leaf image dataset that anchors the comparison table's claim that existing sets are small, single-task, and leaf-only."},{"cited_title":"We also provide additional results in the supple- mentary files","cited_arxiv_id":null,"evidence_quote":"The cotton leaf disease dataset used as another scope baseline and in cross-dataset classification and enhancement runs."},{"cited_title":"This dataset enables exten- sive applications such as classification, segmentation, and restoration by encompassing aerial and close-up images across various growth stages","cited_arxiv_id":null,"evidence_quote":"YOLOWeeds, the cotton-production weed detection benchmark that supplies the YOLO detector framework used for the detection experiments."},{"cited_title":"Deep spectral methods: A surprisingly strong baseline for unsupervised semantic segmentation and localization,","cited_arxiv_id":null,"evidence_quote":"Deep Spectral Method, the unsupervised segmentation baseline applied to the aerial images to demonstrate the segmentation annotations."},{"cited_title":"Blind/referenceless image spatial quality evaluator,","cited_arxiv_id":null,"evidence_quote":"CLIP, the vision-language backbone for the zero-shot and linear-probing classification results that benchmark the dataset's learnability."},{"cited_title":"Clipscore: A reference-free evaluation metric for image captioning,","cited_arxiv_id":null,"evidence_quote":"The VGG19 architecture comparison that the disease classifier is built on, producing the 83.37% test accuracy."},{"cited_title":"Bayesian Image Reconstruction using Deep Generative Models","cited_arxiv_id":"2012.04567","evidence_quote":"Bayesian image reconstruction using generative models (BRGM), one of the two restoration pipelines applied to drone and DSLR splits."},{"cited_title":"Nima: Neural im- age assessment,","cited_arxiv_id":null,"evidence_quote":"Robust unsupervised StyleGAN image restoration (RUSIR), the framework used for denoising, deartifacting, upsampling, and inpainting evaluations."}],"review_version":1}