{"id":"5818765d-81ed-4273-81d5-744dced830a5","arxiv_id":"2505.22065","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AquaMonitor is a 2.7M-image, multi-view, multimodal dataset of aquatic invertebrates collected during routine monitoring, with three benchmark tasks and baseline results.","lead":"Researchers assembled a 2.7-million-image dataset of aquatic invertebrates from two years of routine Finnish lake monitoring, with DNA and biomass data for subsets. It offers a realistic testbed for automated species identification under the imbalances and surprises of actual monitoring programs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'unbiased setup' claim is not supported by the coverage analysis in Sec. A.4.1: 27% of 2022 specimens, 9 taxa, and the majority of 2022 sites are missing from the imaged dataset, making the monitoring benchmark potentially unrepresentative of routine monitoring.","rationale":"The reader's weakest assumption points to the representativeness of the imaged subset, and I find that this is indeed the most load-bearing concern. The abstract and Sec. 3.3 claim an 'unbiased setup', but the supplementary coverage analysis (A.4.1) directly undermines this: only a minority of sites pass the bootstrap representativeness test, and entire lakes and taxa are missing. This is not a minor caveat because the dataset's primary purpose is to serve as a realistic, unbiased evaluation target for monitoring. If the missing specimens are systematically different (as size-based exclusion suggests), benchmark conclusions may not transfer to operational use. The paper already provides the raw numbers, but the main text does not temper the 'unbiased' wording accordingly. My proposed check would quantify the impact on the headline metrics, settling whether the concern is material. I agree with the reader that the verdict should remain conditional: the dataset is valuable and the analysis is honest, but the claim needs revision or additional evidence before it can be accepted as an unbiased setup.","tokens_in":38334,"tokens_out":6225,"duration_ms":70809,"concrete_test":"Using the full monitoring database counts referenced in Sec. A.4.1, reweight the 2022 monitoring test set by site-taxon to match the full 2022 monitoring distribution, then recompute the ensemble model's accuracy and weighted F1 from Table 5. If these reweighted metrics differ by more than 5% from the reported values, the 'unbiased' setup claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that AquaMonitor provides an 'unbiased setup' for evaluating real-life monitoring (Abstract, Sec. 3.3) rests on the assumption that the imaged subset faithfully represents the monitoring program's target distribution. The paper's own coverage analysis (Sec. A.4.1) shows this assumption fails for a large fraction of the data: only 16/41 (2021) and 21/29 (2022) sites are statistically well-represented, nine taxa were not imaged (too large/small for the device), and 6 lakes from 2022 are entirely missing. Since the un-imaged specimens are not a random sample (e.g., size-dependent exclusion, lost containers, site-level non-delivery), the monitoring benchmark's test set (2022) does not mirror the distribution a deployed system would encounter. If the missing specimens are systematically different in size, species composition, or site-specific appearance, the reported accuracy and OOD detection numbers could either overstate or understate real-world performance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AquaMonitor, a large multi-view image-sequence dataset of aquatic macroinvertebrates collected in congruence with an operational Finnish freshwater monitoring program across 2021 and 2022. The dataset contains 2.7M images from 43,189 specimens, with subsets carrying DNA barcodes, dry mass, and size measurements. The authors define three benchmarks: a monitoring benchmark with temporal train/test split, open-set classes, and extreme class imbalance; a standard closed-set classification benchmark; and a few-shot benchmark. They report baselines for all three, plus biomass regression, with bootstrapped confidence intervals. The central claim is that the dataset provides a realistically challenging and unbiased setup for evaluating automated identification methods for routine aquatic biodiversity monitoring.","tokens_in":38574,"tokens_out":2262,"duration_ms":26998,"significance":"If the dataset and benchmarks deliver on their claims, this is a valuable community resource. Its strengths include: standardized kick-sampling and expert taxonomic identification independent of the imaging study; a realistic temporal split that reflects deployment conditions; inclusion of rare and hard-to-identify taxa rather than a curated easy subset; rich metadata (site, time, DNA, biomass) that enables multimodal and ecological analyses; and reproducible baselines with released code and model weights. The multi-view sequence format is also relatively rare among invertebrate datasets. The paper is transparent about many coverage limitations in the supplementary material, which is commendable. However, the headline claim of an 'unbiased setup' is not supported by the paper's own coverage analysis, and this needs to be resolved before the paper can be accepted.","major_comments":[{"comment":"The claim that AquaMonitor provides an 'unbiased setup' (Abstract) and 'represents the full diversity of species encountered during regular biomonitoring, avoiding selection bias' (Sec. 3.3) is contradicted by the coverage analysis in Sec. A.4.1. Only 16/41 sites in 2021 and 21/29 sites in 2022 are statistically well-represented; 9 of 161 taxa were not imaged at all; 6 lakes from 2022 are entirely absent; and the missing 2022 specimens amount to 27.35% of the monitoring-program total. The stated reasons for missingness—specimens too large or too small for the imaging device, lost containers, and site-level non-delivery—are not random with respect to taxon size or site, so the imaged subset is not a random sample of the monitoring distribution. This directly affects the monitoring benchmark's test set (2022), which is supposed to mirror real-life deployment. The authors should either soften the 'unbiased' claim substantially or provide evidence (e.g., comparison of size distributions, taxonomic compositions, or site characteristics between imaged and missing specimens) that the missingness is ignorable for the benchmark conclusions.","section":"Abstract, Sec. 3.3, Sec. A.4.1"},{"comment":"The monitoring benchmark conclusions are based on a test set whose representativeness is compromised by the coverage gaps. In particular, the OOD detection evaluation uses only 72 specimens from 24 OOD classes (Sec. A.3). With such a small outlier set, the reported AUROC values (e.g., 0.80 for the ensemble) have wide confidence intervals and are sensitive to the species composition of the missing 2022 specimens. The paper should discuss how the absence of 6 lakes and 9 taxa in the imaged data affects the difficulty and realism of the OOD and classification results, and should avoid stating that the benchmark 'includes all the challenges encountered in an operational monitoring setting' without qualification.","section":"Sec. 4.1, Table 5, Fig. 4"}],"minor_comments":[{"comment":"The text says 'all but 17 containers from the year 2021' twice; the second occurrence should presumably refer to 2022. Please correct this typo.","section":"Sec. A.4.1"},{"comment":"The phrase 'the dataset remains unbiased' in Sec. A.4.1 is used in a different sense than the abstract's 'unbiased setup.' The supplementary definition (imaging all received specimens without selection) does not imply statistical representativeness. Please use distinct terminology to avoid confusion.","section":"Sec. 3.2.2 and Sec. A.4.1"},{"comment":"The distinction between 'Unique' and 'Labeled to this level' is not immediately clear from the caption; a short explanation in the table caption would help.","section":"Table 3"},{"comment":"The figure lists nine missing taxa, but the caption does not state whether these are the only missing taxa or just a summary; please clarify that this is the complete list referenced in Sec. 3.2.2.","section":"Fig. 9"}],"recommendation":"major_revision","confidential_remarks":"The dataset is potentially a strong contribution, and the authors are transparent about many limitations. The main issue is the mismatch between the 'unbiased' claim and the coverage statistics; this is fixable through revised claims and/or additional analysis. I would encourage the editor to request a revision rather than reject, as the resource and benchmarks are likely to be useful to the community. I also note that the paper's self-citation of BIODISCOVER papers is appropriate given the shared imaging device."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the dataset is a real resource and the baselines are careful. But the abstract and Section 3.3 say 'unbiased', and the supplementary's own coverage analysis says otherwise. That mismatch needs fixing before I'd want to see this pressed as-is.\n\nWhat's new: AquaMonitor is the first aquatic invertebrate image dataset collected in congruence with a routine national monitoring program. That matters, because prior datasets (FINBenthic2, Høye et al., Simović) are smaller, closed-set, or selection-biased. You get 2.7M multi-view images of 43k specimens, with site and time metadata, plus DNA and biomass subsets. The monitoring benchmark is well-designed: train on 2021, test on 2022, 85 ID classes plus 24 OOD classes, and the long tail is genuinely long. Baselines span a reasonable set of architectures, and bootstrapped confidence intervals are the right call. Data and code appear to be public.\n\nWhere it wobbles: the coverage numbers. 89.6% of 2021 specimens and 72.7% of 2022 specimens were imaged. Nine taxa were never imaged because they are too big or too small for the device. Six lakes in 2022 simply don't appear. The site-level representativeness analysis in A.4.1 finds only 16/41 (2021) and 21/29 (2022) sites 'well-represented'. Missingness is not random: it's driven by size limits, lost containers, and logistics. So the 2022 test set is not a faithful sample of what the monitoring program encounters. The authors disclose all this, then in the same section say 'ensuring that the dataset remains unbiased.' That's a step too far. It is a careful and largely honest dataset, but the 'unbiased setup' phrase in the abstract and Section 3.3 overstates what the data can support. The OOD and distribution-shift results in particular could shift if the missing specimens were included.\n\nFor the record: I don't think this is fatal. The benchmark is still the most realistic aquatic invertebrate benchmark available. But the claims need to be softened to 'unbiased with respect to what we could image' or the analysis needs to characterize the missing specimens' size/taxonomy and show the benchmark conclusions are robust to reweighting. The authors clearly have the data to do this.\n\nWho it's for: CV researchers working on fine-grained classification, long-tail, open-set, and multi-view recognition, plus ecologists who want a realistic benchmark. I'd be glad to see this in the literature after revision. It deserves a serious referee; send it to review, but the referee should push hard on the representativeness claim.","headline":"A genuinely useful monitoring-grade aquatic invertebrate dataset with strong baselines, but the 'unbiased' framing doesn't survive its own coverage numbers.","tokens_in":39076,"tokens_out":3410,"would_cite":true,"duration_ms":32927,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AquaMonitor: a 2.7M-image dataset of aquatic invertebrates captured during real routine monitoring, with benchmarks that reflect deployment conditions rather than curated closed-set galleries.","keywords":["AquaMonitor","aquatic invertebrates","biodiversity monitoring","multi-view image sequences","fine-grained classification","few-shot learning","open-set recognition","biomass estimation"],"falsifier":"Take the 2022 test predictions and split them into the 21 statistically representative sites versus the rest, then compare accuracy, macro-F1, and OOD AUROC with bootstrap confidence intervals; a material gap would falsify the unbiased-testbed claim. A second check is to measure the body-size distribution of the nine un-imaged taxa against the imaged taxa and show whether the missing specimens sit at the size extremes.","tokens_in":38194,"feed_emoji":"🐛","tokens_out":5400,"duration_ms":58339,"temperature":0.7,"pith_summary":"The paper introduces AquaMonitor, a dataset built by imaging every feasible specimen from two years of an operational freshwater monitoring program in Finland, yielding 2.7M images from 43,189 specimens and 152 classes. Its central claim is that this is the first large computer-vision dataset of aquatic invertebrates collected in congruence with routine monitoring, so the class distribution, rare taxa, and year-to-year turnover match what a deployed system would actually face. The authors define three benchmarks: a monitoring benchmark that trains on 2021 and tests on 2022 with out-of-distribution classes and extreme imbalance, a standard closed-set classification benchmark, and a few-shot benchmark for rare classes. They also provide DNA barcodes, biomass measurements, size measurements, and site/time metadata for subsets, and they report strong baselines showing that the realistic monitoring task is considerably harder than standard classification. If the unbiased-setup claim holds, progress on the monitoring benchmark can transfer directly to legislative water-quality assessment.","feed_headline":"2.7M aquatic-invertebrate images from routine monitoring debut","feed_subtitle":"Dataset pairs each specimen with multi-view images plus DNA and biomass to test identification under real conditions.","key_machinery":"The load-bearing object is the dataset itself, generated by the BIODISCOVER imaging device, a named multi-view system that drops each specimen through a cuvette while two perpendicular cameras photograph it at 50 frames per second, producing synchronized image sequences. The benchmark design carries the argument: the monitoring benchmark splits by year (2021 train, 2022 test), so the test set naturally contains distribution shift, extreme class imbalance, and classes unseen at training time; this is the setup that makes the dataset a realistic testbed rather than another curated gallery. The additional modalities, COI DNA barcodes, biomass, body-size measurements, and site/time metadata, are the parts that let future evaluations go beyond image-only classification.","core_discovery":"The paper's central claim is that AquaMonitor is the first large computer-vision dataset of aquatic invertebrates assembled in congruence with an operational monitoring program, so its class distribution, rare taxa, and year-to-year turnover reflect what a deployment would actually encounter. Every specimen that could be imaged was imaged: 44,854 synchronized two-camera sequences, 2.7M frames, and 43,189 specimens, with 152 classes spanning a hierarchical taxonomy that includes life-stage variants. On top of the images, subsets carry DNA barcodes (1,358 specimens), dry-mass and size measurements (1,494 specimens), and per-specimen sampling site and time. The authors define a monitoring benchmark (train on 2021, test on 2022, with 24 out-of-distribution classes), a closed-set classification benchmark on 42 well-populated classes, and a few-shot benchmark on 47 rare classes, and they report baselines for each. Their results show that the realistic monitoring benchmark is much harder than standard closed-set classification, with best weighted F1 around 0.859 versus roughly 0.988 on the classification benchmark.","pith_inferences":["Because only 16 of 41 (2021) and 21 of 29 (2022) sampling sites were statistically representative of the monitoring database, a natural next experiment is to rerun the monitoring benchmark on those sites only; if results shift beyond the reported bootstrap intervals, the 'unbiased testbed' claim would need qualification.","The nine missing taxa are described as too large or too small to image, which implies the dataset underweights size extremes; a size-conditioned error analysis would show whether that blind spot affects practical estimates.","DNA barcodes cover only 23 classes, but for those classes they could be used as auxiliary supervision or as a ground-truth signal for out-of-distribution detection, a multimodal fusion the paper lists as future work.","The year-split monitoring protocol could be exported to other routine monitoring programs; if adopted widely, comparable benchmarks across countries would emerge."],"forward_implications":["Performance on the monitoring benchmark is a direct proxy for how automated identification would behave in a real routine freshwater monitoring program, so advances there can feed legislative water-quality assessment.","The large gap between classification-benchmark accuracy (about 0.988) and monitoring-benchmark weighted F1 (about 0.859) shows that closed-set evaluations overstate readiness for deployment.","Out-of-distribution detection with standard ranking scores reaches only about 0.80 AUROC on 72 specimens from 24 unseen classes, marking open-set handling as the current bottleneck.","Using both camera views and averaging logits over the sequence improves performance over single-frame, single-view classification, so the multi-view sequence format carries usable signal.","Biomass regression models initialized from AquaMonitor classification features beat ImageNet-initialized models, so the dataset can support trait estimation as well as identification."],"supporting_citations":[{"why":"Supplies the BIODISCOVER multi-view imaging device and the biomass-estimation methodology that the dataset is built on.","marker":"[3]"},{"why":"Provides the FINBenthic2 aquatic invertebrate dataset, the largest earlier dataset from the same imaging device, used as a comparison point.","marker":"[4]"},{"why":"BIOSCAN-5M is the reference for a self-imaged, image-DNA insect dataset and the main comparison for dataset scale and metadata.","marker":"[24]"},{"why":"Earlier BIODISCOVER-based aquatic invertebrate dataset that is smaller and evaluated closed-set, establishing the gap AquaMonitor fills.","marker":"[36]"},{"why":"AMI dataset from in-situ camera traps, used to contrast operational collection of terrestrial flying insects with the aquatic lab-imaged setup.","marker":"[39]"},{"why":"Recent aquatic invertebrate image dataset with multi-view images, used as a smaller comparison for specimen and class counts.","marker":"[73]"},{"why":"BioCLIP and TreeOfLife-10M supply pretrained species-classification models used as baselines and for transfer in the benchmarks.","marker":"[75]"},{"why":"Provides the fwhF2/FwhR2n COI primers used to sequence the DNA subset.","marker":"[80]"},{"why":"iNat21 represents the citizen-science benchmark context and the kind of curated, charismatic-species data AquaMonitor contrasts with.","marker":"[83]"},{"why":"Documents the MaaMet monitoring program that supplied the specimens and the standardized sampling protocol.","marker":"[88]"}],"fun_headline_variants":["First large aquatic-invertebrate CV dataset from real monitoring","2.7M images, 43K specimens: monitoring-grade invertebrate dataset","Routine water monitoring yields 2.7M-image invertebrate dataset","Aquatic invertebrate dataset with DNA, biomass, and multi-view photos","Realistic benchmarks for aquatic invertebrate ID from 2.7M images"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The impartiality of the dataset depends on the assumption that the specimens that were not imaged (about one in ten in 2021 and more than one in four in 2022, including nine entire taxa) are not systematically different from those that were; if they are, benchmark scores may not predict real monitoring performance.","fun_headline_variants_meta":{"raw":{"variants":["First large aquatic-invertebrate CV dataset from real monitoring","2.7M images, 43K specimens: monitoring-grade invertebrate dataset","Routine water monitoring yields 2.7M-image invertebrate dataset","Aquatic invertebrate dataset with DNA, biomass, and multi-view photos","Realistic benchmarks for aquatic invertebrate ID from 2.7M images"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000225,"raw_usage":{"total_tokens":1497,"prompt_tokens":1012,"completion_tokens":485,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":389}},"tokens_in":628,"tokens_out":485,"duration_ms":5228,"temperature":1.0,"reasoning_tokens":389,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:15:12.426723+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 2022 test predictions and split them into the 21 statistically representative sites versus the rest, then compare accuracy, macro-F1, and OOD AUROC with bootstrap confidence intervals; a material gap would falsify the unbiased-testbed claim. A second check is to measure the body-size distribution of the nine un-imaged taxa against the imaged taxa and show whether the missing specimens sit at the size extremes.","supporting_citations":[{"cited_title":"Gharaee, S","cited_arxiv_id":null,"evidence_quote":"BIOSCAN-5M is the reference for a self-imaged, image-DNA insect dataset and the main comparison for dataset scale and metadata."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier BIODISCOVER-based aquatic invertebrate dataset that is smaller and evaluated closed-set, establishing the gap AquaMonitor fills."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"AMI dataset from in-situ camera traps, used to contrast operational collection of terrestrial flying insects with the aquatic lab-imaged setup."},{"cited_title":"Simovi´c, A","cited_arxiv_id":null,"evidence_quote":"Recent aquatic invertebrate image dataset with multi-view images, used as a smaller comparison for specimen and class counts."},{"cited_title":"Stevens, J","cited_arxiv_id":null,"evidence_quote":"BioCLIP and TreeOfLife-10M supply pretrained species-classification models used as baselines and for transfer in the benchmarks."},{"cited_title":"Vamos, V","cited_arxiv_id":null,"evidence_quote":"Provides the fwhF2/FwhR2n COI primers used to sequence the DNA subset."},{"cited_title":"Van Horn, E","cited_arxiv_id":null,"evidence_quote":"iNat21 represents the citizen-science benchmark context and the kind of curated, charismatic-species data AquaMonitor contrasts with."},{"cited_title":"Vilmi, M","cited_arxiv_id":null,"evidence_quote":"Documents the MaaMet monitoring program that supplied the specimens and the standardized sampling protocol."}],"review_version":1}