{"id":"dbeec1a5-6ed7-48d5-b5f0-2e8897c422d1","arxiv_id":"2504.15252","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"SuoiAI is a proposed end-to-end pipeline for building the first aquatic invertebrate image dataset in Vietnam and using semi-supervised and fine-grained machine learning for species classification.","lead":"This paper describes a planned pipeline to collect images of aquatic invertebrates in Vietnamese streams and train machine learning models to identify them. It is a proposal and roadmap, not a finished dataset or trained model, so a generalist reader should treat it as a work plan rather than a result.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The proposal's annotation budget is internally inconsistent: a 'few hundred' manual labels cannot seed 100-200 species classes at 'about 1000 labeled samples per class,' so the claimed annotation-cost reduction rests on an untested teacher-student bootstrap.","rationale":"Good-faith reading: this is a two-page workshop proposal, not a claim of completed results. As the reader notes, there is no dataset, code, or experiment to verify, so UNVERDICTED is appropriate. My concern picks a point where the proposal can be checked against its own text. The discrepancy between the 'few hundred' manual seed and the '1000 labeled samples per class' target exposes the key unverified mechanism: the teacher-student loop must create reliable fine-grained labels at scale. This is not a disagreement with consensus; it is an internal feasibility question. The proposed test is cheap because public fine-grained datasets exist. If the bootstrap works well, the concern is retired; if it fails, the core value proposition needs rethinking. I do not see fraud or bad faith; the paper is transparent about being a proposal, and its references are relevant. Verdict remains UNVERDICTED, consistent with the reader.","tokens_in":4495,"tokens_out":3908,"duration_ms":36497,"concrete_test":"Run a controlled bootstrap simulation on an existing fine-grained invertebrate or insect dataset (e.g., the Insect Foundation 1M dataset or an iNaturalist subset with 100 classes): train a teacher on 200 randomly selected labeled images, have it pseudo-label 100k unlabeled images, train a student on those pseudo-labels, and compare the student's species-level accuracy against a student trained on 100k human labels. If the pseudo-label student's accuracy is substantially lower, or the pseudo-label noise rate is high, the proposed annotation-cost reduction is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that SuoiAI can build a usable aquatic-invertebrate dataset while reducing annotation effort through semi-supervised learning. The paper's own numbers make this claim insecure. Section 2.2 ('Annotation Strategy') says labeling will begin with 'a few hundred high-quality images for genus and species identification,' and a teacher model will be trained on those images to label additional data. Section 3.3 ('Fine-grained Image Classification') then specifies 20-50 genus-level classes and 100-200 species-level classes 'with about 1000 labeled samples per class.' Reaching 1000 samples per species across 100-200 species requires 100,000-200,000 final labels. If the few hundred manual labels are the total seed, the teacher must bootstrap roughly three orders of magnitude more labels from a few examples per class, and for fine-grained species ID, pseudo-label noise will propagate through the student model. If the few hundred figure is meant per class, the paper should say so, but then the annotation-cost reduction claim is weaker because initial labeling cost is not small. Either way, the load-bearing assumption is an untested teacher-student bootstrap that can turn a tiny seed into accurate species-level labels. The paper honestly describes deployment robustness as a requirement rather than a demonstrated capability; I agree with the reader that camera-image quality is a real risk, but the labeling-budget tension is more internal and checkable from the text alone.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SuoiAI, an end-to-end pipeline for collecting and labeling underwater images of aquatic invertebrates in Vietnam and for training object detection and fine-grained classification models on them. It describes field deployment in two national parks, a semi-supervised annotation strategy using a teacher-student bootstrap, and several modeling approaches. The paper is explicitly a proposal: it reports no dataset, no experimental results, and no trained models.","tokens_in":4880,"tokens_out":7734,"duration_ms":62520,"significance":"The problem is important: Vietnam's freshwater invertebrate fauna is under-documented, and a publicly available dataset would be a valuable resource for ecology and conservation. If the proposed pipeline works as described, it could substantially reduce the annotation cost of building such a dataset and serve as a template for other tropical regions. The paper is transparent about being a proposal, and it draws on a reasonable set of established techniques (SAM, CLIP, semi-supervised object detection, fine-grained classification, clustering-based annotation). The main concern is that the central claim of annotation-cost reduction rests on an internally inconsistent budget, and none of the pipeline components is demonstrated on any data.","major_comments":[{"comment":"The annotation strategy in §2.2 begins with manual labeling of 'a few hundred high-quality images' for genus and species identification, while §3.3 specifies a target of 100–200 species-level classes 'with about 1000 labeled samples per class.' If these are the same labels, the seed set is three orders of magnitude smaller than the final budget (100,000–200,000 labels), requiring the teacher model to bootstrap almost the entire dataset; if the 'few hundred' is meant per class, then the total manual effort is tens of thousands of labels, which contradicts the paper's emphasis on reducing annotation cost. The paper does not provide a feasibility analysis for either reading, so the cost-reduction claim is not supported.","section":"§2.2 and §3.3"},{"comment":"The title 'Building a Dataset' and parts of the abstract imply that a dataset has been or is being built, but no empirical evidence of any kind is presented: there is no pilot dataset, no evaluation of the teacher-student semi-supervised loop, no detection or classification accuracy, and no field data from the two mentioned national parks. Without at least a proof-of-concept on a small number of classes, the reader cannot assess whether the proposed pipeline is viable, and the paper's contribution is difficult to evaluate. The authors should either add a small-scale validation or explicitly reframe the paper as a project plan.","section":"Abstract and Title"},{"comment":"The paper acknowledges that 'the system must be robust in diverse aquatic conditions, addressing variables such as lighting, turbidity, and background noise,' but provides no evidence that the proposed underwater cameras will produce usable images at the planned sites, which is a prerequisite for the entire pipeline. A small pilot deployment at Cat Tien or Cuc Phuong, reporting image quality and specimen visibility, would strengthen the proposal; without it, the feasibility of data collection is an untested assumption.","section":"§4 Practical Considerations"}],"minor_comments":[{"comment":"The phrase 'Suoi means streams in Vietnamese' would be more precise as 'Suối (Suoi) means stream in Vietnamese'; the plural form is not needed in the name.","section":"§1 (footnote on 'Suoi')"},{"comment":"The production estimate of 'around three million data points per site annually' is asserted without derivation; the authors should specify the assumed frame rate, active hours, and image size.","section":"§2.1"},{"comment":"The sentence 'We will also look to include existing datasets of regional invertebrates to create the initial dataset' is vague; please name the datasets and describe how class overlap with new images will be handled.","section":"§2.2"},{"comment":"The target of '100-200 species-level classes' with 'about 1000 labeled samples per class' is presented without discussion of class imbalance or how rare species will be treated in the long-tail distribution; this is relevant to the evaluation plan.","section":"§3.3"},{"comment":"Figure 1 is not referenced in the text; the reader is not told where in the pipeline to look at it, and the individual components in the diagram are not explained in a caption.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop paper and is transparent about being a proposal. For a standard journal, the absence of any evaluated component is a serious concern, and the internal inconsistency in the annotation budget accentuates it. I would ask the authors to provide a pilot study or to reframe the paper explicitly as a project plan, and to resolve the annotation-budget arithmetic. The reference list is appropriate for the proposal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou can read this one in ten minutes and then set it aside. It's a two-page workshop proposal (ICLR 2025 climate workshop) for building a dataset of aquatic invertebrates in Vietnamese streams. There is no dataset, no model, no evaluation. The reader's 'UNVERDICTED' is right: there is no claimed result to check.\n\nWhat is actually new is narrow but real: a focused pitch for closing a documented gap in biodiversity data for Vietnam, using underwater cameras plus semi-supervised labeling. The authors name relevant existing datasets and methods, and they are honest that deployment robustness is an open problem. As a proposal, the writing is clear and the pipeline is plausible. The citation pattern looks fine.\n\nThe soft spot is the annotation budget. Section 2.2 says manual labeling starts with 'a few hundred high-quality images'; Section 3.3 then asks for 100–200 species classes at 'about 1000 labeled samples per class,' i.e. 100k–200k labels. The stress-test note is right that these numbers are not reconciled. Maybe the 'few hundred' is just the seed and the teacher-student bootstrap is expected to generate the rest, or maybe existing datasets will help — but the paper doesn't say. For a proposal, this is a fixable clarity problem, not a fatal one, because no results are claimed. Still, if the authors want the cost-reduction story to carry weight, they should give rough counts or a pilot plan.\n\nWho is this for? People working on ML for biodiversity in data-scarce tropical regions, and workshop organizers looking for proposal-style papers. It does not deserve a full journal referee cycle in its current form, because there is no artifact to review. For a workshop, it is fine. For a more formal venue, it should be returned with an invitation to resubmit after collecting even a small pilot dataset.\n\nMy recommendation: desk-reject for a results-oriented venue; accept the proposal framing only in a venue that explicitly takes proposals. I would not cite it, but I would be interested to see the dataset if they build it.","headline":"Honest two-page workshop proposal for an aquatic-invertebrate dataset pipeline in Vietnam; no results to verify, and the annotation-budget numbers don't add up, but the gap it targets is real.","tokens_in":5285,"tokens_out":2340,"would_cite":false,"duration_ms":19867,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SuoiAI proposes that an end-to-end pipeline of underwater cameras, semi-supervised labeling, and machine learning can build Vietnam's first aquatic-invertebrate dataset and classify species from it.","keywords":["aquatic invertebrates","Vietnam biodiversity","underwater camera monitoring","semi-supervised learning","teacher-student pseudo-labeling","fine-grained species classification","dataset construction pipeline","freshwater ecology"],"falsifier":"Deploy the proposed cameras at Cat Tien and Cuc Phuong through a full wet-dry cycle, then measure the fraction of frames in which a freshwater biologist can identify at least one invertebrate. If that fraction is too low to collect a few hundred high-quality manually labeled images within a season, the pipeline's central feasibility claim fails.","tokens_in":4274,"feed_emoji":"🌊","tokens_out":9191,"duration_ms":75763,"temperature":0.7,"pith_summary":"SuoiAI is a proposed pipeline that would build a labeled image dataset of aquatic invertebrates in Vietnam and use machine learning to classify the species. The paper argues that a few hundred expert-labeled images, combined with underwater cameras and a loop in which an initial model labels new images and later models learn from those labels, can grow into thousands of labeled samples without a large manual annotation budget. This matters because Vietnam is severely underdocumented: of more than 100,000 described freshwater invertebrate species worldwide, only about 2,000 have been identified in Vietnam, and 800 aquatic species lack systematic documentation. If the pipeline works, ecologists gain a scalable tool for monitoring water quality and biodiversity in a climate-sensitive tropical region. The paper is explicitly a proposal; it reports no dataset, model, or experimental results yet.","feed_headline":"Vietnam's stream invertebrates get an AI dataset pipeline","feed_subtitle":"Underwater cameras plus semi-supervised labeling could turn a 2,000-species record into a scalable monitoring tool.","key_machinery":"The load-bearing mechanism is the teacher-student bootstrap labeling loop, assisted by clustering and prompt-based segmentation. A teacher model trained on a few hundred expert-labeled images pseudo-labels new unlabeled captures; student models are trained on the mixture, and the cycle repeats. MorphoCluster groups visually similar images so annotators label clusters rather than single images, and SAM produces segmentation masks from prompts without training. This loop is what turns a small human effort into the large, fine-grained dataset the classification stage requires.","core_discovery":"The claim, stated on the paper's own terms, is that SuoiAI can deliver what has been missing for Vietnamese stream ecology: a labeled image dataset large enough and fine-grained enough to train species-level classifiers for aquatic invertebrates. The envisioned pipeline couples automated 1080p underwater cameras at two national parks with a few hundred expert-labeled images, then expands the labels through a teacher-student semi-supervised loop assisted by clustering and prompt-based segmentation. The paper expects roughly 20-50 genus-level classes and 100-200 species-level classes, with about 1,000 labeled samples per class, growing to 135 sites and about three million data points per site per year. If that holds, species classification and population tracking for Vietnamese stream invertebrates become feasible without a large manual annotation budget.","pith_inferences":["Beyond the paper's explicit claims, the transferable core is the bootstrap annotation loop, not any particular camera or detector; the same teacher-student plus clustering pattern applies to other underdocumented taxa where images are cheap and labels expensive.","Because the design targets 10-50 mm specimens with 1-5 individuals per frame, it is implicitly a benthic-macroinvertebrate system; applying it to planktonic or swimming fauna would require different camera placement and capture rates, not just different models.","A decisive early test would be a per-class annotation-cost comparison: if pseudo-labeling does not reduce effort on rare species relative to manual labeling alone, the economic argument collapses even if image quality is adequate."],"forward_implications":["Vietnam would obtain its first systematic aquatic-invertebrate image dataset, shrinking the gap between the roughly 2,000 documented species and the estimated diversity.","Annotation cost per usable image drops enough that biodiversity datasets become feasible in regions that cannot hire large labeling teams.","Repeated deployments at 135 sites would allow population and range tracking for stream invertebrates, giving conservation managers a direct signal for water quality and climate effects.","A pretrained aquatic-invertebrate foundation model could emerge from the curated labels, letting downstream users classify images from nearby tropical regions with little or no extra training."],"supporting_citations":[{"why":"Supplies the data-gap numbers, roughly 2,000 identified freshwater invertebrate species in Vietnam and 800 lacking documentation, that motivate the dataset.","marker":"Thuaire et al. (2021)"},{"why":"Presents the iNaturalist classification and detection dataset, the main existing biodiversity image dataset the pipeline positions itself against.","marker":"Horn et al. (2018)"},{"why":"Provides the mixed pseudo-labels teacher-student semi-supervised object detection method the pipeline adopts.","marker":"Chen et al. (2023)"},{"why":"Introduces SAM 2, used for prompt-based segmentation to create masks for annotation without training.","marker":"Ravi et al. (2024)"},{"why":"Describes MorphoCluster, the clustering approach used to group similar images and streamline targeted annotation.","marker":"Schröder et al. (2020)"},{"why":"Offers deep hierarchical Bayesian learning for classifying unknown insect species, referenced for handling rare or unknown taxa.","marker":"Badirli et al. (2023)"},{"why":"Presents PEEB, a part-based fine-grained image classifier, as the fine-grained classification technique the pipeline plans to use.","marker":"Pham et al. (2024)"},{"why":"Demonstrates zero-shot insect detection via weak language supervision, supporting classification of unseen species.","marker":"Feuer et al. (2024)"},{"why":"Shows few-shot underwater fish species classification with limited data, supporting the few-shot route for scarce labels.","marker":"Villon et al. (2021)"},{"why":"Provides Faster R-CNN, named as a deployable baseline object detector for the pipeline.","marker":"Ren et al. (2016)"}],"fun_headline_variants":["AI pipeline builds Vietnam's aquatic invertebrate dataset","Semi-supervised AI labels stream bugs in Vietnam","Underwater cameras + AI to catalogue Vietnam's stream life","New AI dataset for Vietnam's stream invertebrates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline stands or falls on whether underwater cameras in Vietnamese streams capture images in which a person or model can actually see the invertebrates; if the water is too cloudy, dark, or cluttered, no labeling or learning method can recover the dataset.","fun_headline_variants_meta":{"raw":{"variants":["AI pipeline builds Vietnam's aquatic invertebrate dataset","Semi-supervised AI labels stream bugs in Vietnam","Underwater cameras + AI to catalogue Vietnam's stream life","New AI dataset for Vietnam's stream invertebrates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1212,"prompt_tokens":776,"completion_tokens":436,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":392,"completion_tokens_details":{"reasoning_tokens":375}},"tokens_in":392,"tokens_out":436,"duration_ms":4159,"temperature":1.0,"reasoning_tokens":375,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:27:56.883292+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Deploy the proposed cameras at Cat Tien and Cuc Phuong through a full wet-dry cycle, then measure the fraction of frames in which a freshwater biologist can identify at least one invertebrate. If that fraction is too low to collect a few hundred high-quality manually labeled images within a season, the pipeline's central feasibility claim fails.","supporting_citations":[{"cited_title":"Classifying the unknown: Insect identification with deep hierarchical bayesian learning","cited_arxiv_id":null,"evidence_quote":"Offers deep hierarchical Bayesian learning for classifying unknown insect species, referenced for handling rare or unknown taxa."},{"cited_title":"Singh, Soumik Sarkar, Nirav Merchant, Arti Singh, Baskar Ganapathysubramanian, and Chinmay Hegde","cited_arxiv_id":null,"evidence_quote":"Demonstrates zero-shot insect detection via weak language supervision, supporting classification of unseen species."}],"review_version":1}