{"id":"7d2ca669-eea8-4410-8049-ea29adf6d7bd","arxiv_id":"2505.06637","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The preprint outlines an AI-based salmon monitoring framework that combines vision-language models, sonar and video fusion, and expert-in-the-loop validation, but reports no results.","lead":"This paper describes a proposed project to use AI that combines video and sonar data with human expert review to count and identify wild salmon in remote Indigenous rivers. It lays out a plan and evaluation criteria, but presents no experimental results.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim of reduced manual effort rests on unmeasured VLM refinement; Fig. 3 shows an uncorrected misclassification and §7 admits performance uncertainty.","rationale":"The reader's verdict of UNVERDICTED is appropriate because the paper makes no measured claims. My stress-test identifies the same load-bearing assumption: the efficiency gain depends on VLM accuracy and refinement being sufficient. I looked for a more technical internal inconsistency—e.g., a contradiction between the Figure 3 failure and the claim that VLM refinement reduces workload—and found one in substance: the paper's own evidence (Figure 3) and its risks section (Section 7) admit the uncertainty, yet the abstract and implementation plan assert the benefit as fact. This is a correctness risk, not a circularity or reproducibility issue: the argument's weakest link is unmeasured VLM performance and an unspecified refinement mechanism. The proposed concrete test would settle whether the concern lands by measuring VLM accuracy and expert workload. Since no evidence exists either way, the verdict remains UNVERDICTED; no change to the reader's assessment.","tokens_in":12847,"tokens_out":3003,"duration_ms":29507,"concrete_test":"Retrospective pilot on existing Yakoun River video: select a balanced set of low-confidence detections from the current YOLO/RT-DETR baseline, have the proposed VLM classify them, and compare against expert labels. Measure (a) VLM top-1 accuracy on low-confidence frames against expert labels, (b) the fraction of VLM outputs that would still be flagged for expert review, and (c) total expert time in the VLM-routed workflow versus a full-review baseline. If VLM accuracy is not above a pre-specified threshold (e.g., 90%) or expert workload does not decrease, the central claim of reduced manual effort is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central promise—automated species identification, counting, and length measurement with reduced manual effort and improved accuracy—depends on the assumption in §3.1 and §4.2 that a pre-trained VLM, after expert-in-the-loop refinement, will classify low-confidence frames accurately enough to reduce expert review workload. No measurements support this. The only concrete evidence, Figure 3, shows an off-the-shelf VLM (OpenAI o1) misclassifying a sockeye as a Chinook, requiring expert correction. Section 4.2 asserts that as expert-reviewed frames accumulate the VLM 'progressively improves' and reliance on manual verification decreases, but gives no refinement mechanism, no error rates, no workload measurements, and no evaluation of downstream detection/counting. Section 7 (Assumptions and Risks) explicitly states performance 'remains uncertain' and that only 'preliminary results' from Yakoun River exist, without numbers. If VLM accuracy on low-confidence frames is too low, routing those frames to a VLM adds an extra layer that may not reduce (and could increase) expert review, collapsing the efficiency gain. The paper is a proposal, not a demonstration; the load-bearing efficiency assumption is unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes an expert-in-the-loop multimodal foundation AI framework for wild salmon monitoring in Indigenous rivers of the Pacific Northwest. It integrates underwater video from counting weirs and sonar units for automated species identification, counting, and length measurement. The proposed pipeline routes low-confidence video frames to a vision-language model (VLM) for refinement, with expert review for remaining uncertain cases, and adapts foundation models such as SAM2 for sonar processing with echogram fusion. The manuscript includes problem motivation, alignment with Sustainable Development Goals, a system architecture, planned evaluation metrics, expected impacts, risk discussion, and a project team description. No experimental results are reported; Sections 5 and 6 present evaluation criteria and expected outcomes, not measurements.","tokens_in":12965,"tokens_out":4475,"duration_ms":48400,"significance":"If the proposed framework works as described, it would address a genuine and pressing need: monitoring data limitations for salmon stocks in remote, roadless rivers, with direct benefits for Indigenous-led conservation and adaptive fisheries management. The interdisciplinary team and existing relationships with Indigenous communities, as well as the commitment to open data and culturally informed co-development, are notable strengths that give the proposal practical credibility. However, the paper's central claims of reduced manual effort, expedited results, and improved accuracy are not validated by any quantitative evidence. The only concrete data point (Figure 3) shows an off-the-shelf VLM misclassifying a sockeye as a Chinook, which illustrates the need for expert review but does not demonstrate that the proposed refinement loop will be accurate or efficient. The paper is a project proposal rather than a completed research study; its value depends on future implementation and evaluation.","major_comments":[{"comment":"The abstract claims 'reducing manual effort, expediting delivery of results, and improving decision-making accuracy,' but the manuscript provides no experimental evidence for any of these outcomes. Section 5 only lists planned evaluation criteria, and Section 7 explicitly states that performance 'remains uncertain' and mentions only 'preliminary results from the Yakoun River' without reporting numbers. Figure 3, the sole empirical illustration, shows a VLM error that requires expert correction. This is a load-bearing gap: if VLM accuracy on low-confidence frames remains low, routing frames through a VLM could add an extra review layer rather than reduce expert workload. The authors must either provide quantitative results from their deployments or substantially reframe the paper as a proposal and temper the achievement-oriented claims.","section":"Abstract and Section 4.2"},{"comment":"The statement that 'as expert-reviewed frames accumulate, the refined VLM progressively improves its performance, reducing reliance on manual verification over time' is an unsupported assertion. The manuscript does not specify how expert corrections are incorporated (e.g., fine-tuning, in-context learning, or retrieval augmentation), nor does it provide learning curves, error-rate measurements, or workload statistics. Without a concrete refinement mechanism and evidence of its effect, the claimed efficiency gain is speculative and cannot be assessed by the reader.","section":"Section 4.2, Paragraph 3"},{"comment":"The sonar processing pipeline, which adapts SAM2 with echogram fusion and uses 'attention-based feature rearrangement' for length measurement, is described only at a high level. No results are reported for detection, tracking, counting, or length estimation, despite Section 5.2 defining appropriate metrics. Since the paper promises a full pipeline including length measurement, this missing validation is central to the claimed contribution. The authors should report at least preliminary results on a defined sonar dataset, or clearly mark these components as untested future work.","section":"Section 4.3 and Figure 8"}],"minor_comments":[{"comment":"The text 'Vison Language Model' appears twice in Figure 3 captions; this should be 'Vision Language Model' or, more precisely, 'vision-language model.'","section":"Figure 3 captions"},{"comment":"The model name 'LLaV A' has a formatting issue; it should be 'LLaVA.'","section":"Section 3.1"},{"comment":"The phrase 'preliminary results from the Yakoun River suggest that automated model analysis is not only feasible but also critical' is vague: no data, metrics, or study details are provided. Please either include the relevant results in a data section or remove the unverifiable claim.","section":"Section 7"},{"comment":"The 'Project Team Description' section is unusual for a research paper and reads like grant-proposal material. Consider moving this content to supplementary material or an acknowledgment-style appendix if the venue permits.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This manuscript reads more like a funded project prospectus than a completed research paper. The lack of any evaluated system means it would not meet the bar for most empirical AI venues. However, the problem is important and the collaborative team is well positioned to produce meaningful results. If the editor is willing to consider the paper as a vision/position paper, the authors must significantly tone down their claims and add a clear statement that no evaluation has been completed. If the journal expects empirical validation, rejection may be more appropriate unless the authors can quickly supply results from their Yakoun River work. I recommend the editor clarify the paper's intended contribution type during the decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a project plan, not a research paper. There are no experiments, no data, and no code. What you get instead is a well-structured proposal for combining multimodal foundation models (VLM, SAM2, CLIP) with expert-in-the-loop review for salmon monitoring in Indigenous rivers. The writing is clear, the team is real, and the ethical framing around Indigenous data sovereignty is thoughtful. The figures illustrating the failure case of an off-the-shelf VLM (Figure 3) are honest and useful: they show exactly why expert review is needed, even if they don't show any improvement from the proposed pipeline.\n\nThe genuinely new elements are small: routing low-confidence frames to a VLM for explanation and verification, fusing echograms with sonar frames, and using expert comments as auxiliary input. These are reasonable ideas worth testing, but they are presented as plans, not results. The evaluation section lists metrics (mAP, MOTA, HOTA, etc.) but no measurements. Section 7 openly admits that model performance 'remains uncertain' and that only 'preliminary results' exist without numbers. The paper's main promise—reduced manual effort via a progressively improving VLM—rests on an untested assumption. If VLM accuracy on low-confidence frames is poor, the added layer could increase expert workload rather than decrease it. The stress-test note is accurate on this point.\n\nI agree with the reader's verdict of UNVERDICTED. There is no circularity, and the lineage to prior work (SALINA, Atlas et al.) is legitimate. The citations look appropriate, and the problem is real. But as a contribution to the literature, this is a proposal. A serious peer-reviewed venue for research results should not accept it in this form; it would need at least one dataset, a baseline comparison, and measured end-to-end performance. It might fit a workshop or a non-archival track, or a funding progress report.\n\nIf you are working on computing for conservation, it's worth reading for the problem framing and the collaboration structure, but I wouldn't cite it yet. I'd send it back to the authors with encouragement: the plan is sound, the partnerships are valuable, and the questions matter. The missing piece is doing the work and reporting numbers.","headline":"A clearly written project overview with no experimental results; the central efficiency claim hangs on an unmeasured VLM-routing assumption.","tokens_in":13600,"tokens_out":1380,"would_cite":false,"duration_ms":15356,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes that multimodal foundation AI, kept in check by fisheries experts, can automate salmon species identification, counting, and length measurement at remote river monitoring sites.","keywords":["multimodal foundation AI","vision-language models","expert-in-the-loop","salmon monitoring","sonar-based fish counting","fisheries management","Indigenous data sovereignty","edge-cloud deployment"],"falsifier":"Run the proposed pipeline for a full field season and compare, against expert labels, the vision-language model's species-classification accuracy on low-confidence frames, the total expert review hours per fish, and the mAP/F1 and MOTA/HOTA numbers against the YOLO/RT-DETR and CFC baselines named in the paper; if the VLM does not beat the base model on the low-confidence subset, or if expert hours do not decline while accuracy holds, the central efficiency claim is refuted.","tokens_in":12562,"feed_emoji":"🐟","tokens_out":8401,"duration_ms":78576,"temperature":0.7,"pith_summary":"This paper proposes that multimodal foundation AI—particularly vision-language models that can describe images in words—can be adapted, with fisheries biologists kept in the loop, to automate the identification, counting, and length measurement of wild salmon at remote river weirs and sonar sites. The authors build on existing video and sonar monitoring systems and argue that routing low-confidence detections through a vision-language model, and only sending uncertain cases to human experts, will cut manual review effort while improving accuracy. Sonar analysis would be upgraded by fusing sonar frames with echograms in a foundation model whose outputs feed tracking, counting, and length estimates. The paper is an architecture and implementation plan: it specifies workflows, evaluation metrics, and a co-development process with Indigenous stewardship partners, and it acknowledges in its assumptions section that performance across sites remains uncertain. It reports no field measurements yet, but if the plan works, monitoring sites that currently depend on preseason forecasts and labor-intensive frame review could shift to adaptive in-season decisions.","feed_headline":"AI and human experts could automate salmon counting in remote rivers","feed_subtitle":"A pipeline sends uncertain fish frames to vision-language models and fisheries biologists for review.","key_machinery":"The load-bearing mechanism is the expert-in-the-loop verification loop: a lightweight detector/classifier on site flags low-confidence frames, a vision-language model (a model trained to answer questions about images in natural language) classifies them with explanations, and only uncertain cases are escalated to fisheries experts, whose corrections update the model. A second mechanism is multimodal sonar fusion: sonar frames and echograms (time-versus-height sonar returns) are encoded separately and merged through attention, with expert comments included as text, inside an adapted SAM2 foundation model for segmentation, tracking, counting, and length measurement. The design's core promise is that expert effort concentrates on the hardest cases while the model automates the rest, so accuracy rises as corrections accumulate.","core_discovery":"On its own terms, this paper's central claim is that a monitoring pipeline combining a base detector/classifier, a vision-language model, and expert review can make salmon monitoring accurate enough and cheap enough for in-season fisheries management in Indigenous rivers. For video, low-confidence detections from single-modality detectors are sent to a vision-language model, which produces an identification plus a textual explanation; if uncertainty remains, the frame and explanation are flagged for a fisheries biologist, and corrections are fed back as training data. For sonar, the paper proposes adapting the SAM2 foundation model to fuse sonar frames and echograms through separate encoders and attention-based fusion, with expert comments as optional text inputs, supporting detection, tracking, counting, and length measurement. The paper explicitly acknowledges in its Figure 3 that an off-the-shelf vision-language model misidentified a sockeye as a Chinook, and uses that as evidence that expert validation is necessary rather than optional. The claim is advanced as a project design: the expected result is stated, but no measurements are reported that would confirm it.","pith_inferences":["A test the paper leaves implicit: measure expert review hours per thousand fish before and after VLM routing is added; if the VLM is not clearly better than the base model on low-confidence frames, the loop could add workload rather than remove it.","The paper's own Figure 3 hints that spawning-phase and rare-species identifications stay hard for off-the-shelf vision-language models; a plausible outcome is that expert workload drops for common species but stays high for the rare, data-poor cases that matter most to conservation.","The evaluation metrics listed in the paper do not set a numeric threshold for what fisheries management standards require; until such a threshold is defined, a reader cannot tell what accuracy would justify a management decision.","The claim that VLM-assisted annotation reduces errors by inexperienced annotators can be tested directly by comparing inter-annotator agreement with and without VLM-generated suggestions."],"forward_implications":["If the vision-language model reduces the fraction of frames needing expert review, weir-based monitoring shifts from semi-automated to largely AI-driven, lowering the manual-labour bottleneck that currently limits in-season data delivery.","Sonar-based monitoring can extend counting and length measurement across the full river width without building weirs, opening data collection for sites that currently lack infrastructure.","Synchronizing sonar and video where both are available should improve tracking and counting through early fusion, compared with either modality alone.","Open-sourcing datasets and models, together with federated learning, could spread the system across multiple Indigenous territories while keeping raw data under community control.","Reliable in-season abundance estimates would let managers respond to actual returns rather than relying on preseason forecasts, supporting selective harvest of healthy stocks."],"supporting_citations":[{"why":"Supplies the prior video-based weir detection and tracking system that this project extends.","marker":"[Atlas et al., 2023]"},{"why":"Establishes the real-time sonar analytics system and edge deployment context that the proposal builds on.","marker":"[Xu et al., 2024]"},{"why":"Provides the CFC sonar dataset and benchmark motivating detection, tracking, and counting in low-signal underwater settings.","marker":"[Kay et al., 2022]"},{"why":"Supplies the LLaVA vision-language model used for low-confidence classification and explanation.","marker":"[Liu et al., 2024a]"},{"why":"Provides the SAM2 foundation model that the paper proposes to adapt for sonar detection and segmentation.","marker":"[Ravi et al., 2024]"},{"why":"Supplies CLIP, the multimodal encoder used to tokenize sonar frames and echograms.","marker":"[Radford et al., 2021]"},{"why":"Supplies the DeepSORT tracking algorithm used to associate salmon across frames.","marker":"[Wojke et al., 2017]"},{"why":"Single-modality detector baseline that the VLM-enhanced pipeline must beat.","marker":"[Wang et al., 2024]"},{"why":"Single-modality detector baseline compared in both video and sonar evaluations.","marker":"[Zhao et al., 2024]"}],"fun_headline_variants":["AI and experts team up to count salmon in remote rivers","Multimodal AI plus expert review for wild salmon counting","Vision-language model and biologists join forces for salmon monitoring","Automated salmon counting with AI and human-in-the-loop","New pipeline uses AI and expert feedback for sustainable salmon fisheries"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The plan assumes that a pre-trained vision-language model, refined through expert feedback, will classify salmon species accurately enough—and the routing will be selective enough—that the number of frames needing expert review actually falls instead of staying about the same.","fun_headline_variants_meta":{"raw":{"variants":["AI and experts team up to count salmon in remote rivers","Multimodal AI plus expert review for wild salmon counting","Vision-language model and biologists join forces for salmon monitoring","Automated salmon counting with AI and human-in-the-loop","New pipeline uses AI and expert feedback for sustainable salmon fisheries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1283,"prompt_tokens":920,"completion_tokens":363,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":283}},"tokens_in":536,"tokens_out":363,"duration_ms":4055,"temperature":1.0,"reasoning_tokens":283,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:36:44.915201+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the proposed pipeline for a full field season and compare, against expert labels, the vision-language model's species-classification accuracy on low-confidence frames, the total expert review hours per fish, and the mAP/F1 and MOTA/HOTA numbers against the YOLO/RT-DETR and CFC baselines named in the paper; if the VLM does not beat the base model on the low-confidence subset, or if expert hours do not decline while accuracy holds, the central efficiency claim is refuted.","supporting_citations":[{"cited_title":"Wild salmon enumeration and monitoring using deep learning empowered detection and tracking","cited_arxiv_id":null,"evidence_quote":"Supplies the prior video-based weir detection and tracking system that this project extends."},{"cited_title":"The caltech fish counting dataset: a benchmark for multiple-object tracking and counting","cited_arxiv_id":null,"evidence_quote":"Provides the CFC sonar dataset and benchmark motivating detection, tracking, and counting in low-signal underwater settings."},{"cited_title":"Simple online and realtime tracking with a deep association metric","cited_arxiv_id":null,"evidence_quote":"Supplies the DeepSORT tracking algorithm used to associate salmon across frames."}],"review_version":1}