Pith. sign in

REVIEW 4 major objections 5 minor 64 references

An automated vision pipeline can estimate tuna purse-seine catch composition from electronic monitoring video with a mean average error of 4.5%, despite the fact that human experts often cannot distinguish bigeye from yellowfin tuna in such

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 06:44 UTC pith:T35WJ6SB

load-bearing objection Solid applied pipeline with a genuinely useful ground-truth protocol, but the headline 4.5% MAE is cherry-picked from post-hoc per-trip model selection and the abstract overstates hierarchical classification. the 4 major comments →

arxiv 2511.15468 v1 pith:T35WJ6SB submitted 2025-11-19 cs.CV

Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners

classification cs.CV
keywords tunapurse seineelectronic monitoringcatch compositioncomputer visioninstance segmentationspecies classificationhierarchical classification
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper claims that a multi-stage computer vision pipeline, combining a YOLOv9 detector, a SAM2 segmenter, ByteTrack tracking, and a hierarchical species classifier, can automatically estimate the species composition of tropical tuna purse-seine catches from onboard electronic monitoring video. The headline result is that 84.8% of individuals are both segmented and classified, with a mean average error of 4.5% when tested on artificial fishing operations recorded with a global-shutter camera. The work is motivated by the need to replace or support human analysts who must review massive amounts of EM footage, and by the finding that nine tuna taxonomy experts agree on only 42.9% of bigeye and 57.1% of yellowfin identifications from images alone. The paper argues that its onboard-verified ground truth—where scientists identify each fish by hand on the conveyor belt—makes the training data and final estimates more reliable than human image-based review.

Core claim

The central claim is that a cascade of off-the-shelf and fine-tuned deep learning components can close the gap between laboratory-style fish identification and operational fisheries monitoring. The authors show that segmenting all fish into a single generic class, then classifying each tracked individual through a sequence of binary decisions (target vs. non-target, skipjack vs. other tunas, bigeye vs. yellowfin), yields better generalization on real fishing operation data than a standard multiclass classifier. They further show that the choice of camera hardware matters dramatically: a rolling-shutter sensor produces a 12.3% mean absolute error in species composition, while a global-shutter

What carries the argument

The pipeline is the central mechanism. A fine-tuned YOLOv9 detector generates bounding-box prompts for each visible fish; SAM2, a foundation model for image segmentation, converts those prompts into precise segmentation masks without fine-tuning. ByteTrack then tracks each individual across frames, so fish are counted once and multiple views of the same fish are pooled. Finally, a hierarchical classifier (RegNetX400 backbone) makes three binary decisions and averages its per-frame predictions per tracked fish, producing a species label. This decomposition separates the segmentation problem from the species-identification problem, allowing the segmentation stage to exploit the generic object-

Load-bearing premise

The artificial fishing operations—where fish are pulled from the catch, identified on deck, and placed back on the conveyor belt—are representative enough of real haul conditions (stacking, occlusion, motion, lighting) that the 4.5% mean average error carries over to actual fishing operations.

What would settle it

Run the trained YOLOv9-SAM2-hierarchical pipeline on real purse-seine sets where every individual is counted and identified by an onboard observer as the fish are transferred to the hold, and compare per-set species percentages. If the mean absolute error on those real sets exceeds, say, 10%—or if it is no better than the 12.3% error reported for the rolling-shutter AFOs—the claim that the pipeline can accurately estimate catch composition under operational conditions is falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the accuracy holds under real conditions, electronic monitoring footage from tuna purse seiners can be automatically converted into per-set species composition estimates, reducing the need for human video analysts and potentially making EM-based monitoring more consistent than observer-based sampling.
  • The hierarchical classification design—first separating easily distinguishable groups before finer distinctions—can be transferred to other fisheries where species are visually similar but differ in management-relevant traits.
  • The large difference between rolling-shutter and global-shutter cameras implies that hardware selection is as important as model architecture for vision-based catch monitoring, and that future deployments should favor global-shutter sensors.
  • The expert-disagreement result suggests that any image-only labeled dataset for bigeye and yellowfin tuna is inherently noisy; future work should either use in-situ verification or model label uncertainty explicitly.
  • Because the pipeline already records depth information from stereoscopic cameras, adding volumetric size and weight estimation is a natural next step that could further enhance stock-assessment data.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The 4.5% error was measured on artificial fishing operations where fish were placed on the belt in a controlled manner; real hauls have higher fish density, more occlusion, and unsteady belt motion, so operational error is likely higher than reported.
  • The ground truth labels come from experts handling fish on board, but those experts may still have subtle biases or occasional errors; since the model learns from them, the model's ceiling is bounded by the reliability of that in-situ identification.
  • Because SAM2 is used without fine-tuning, the segmentation stage could likely be swapped for a newer foundation model without retraining the rest of the pipeline—an easy testable extension as these models improve.
  • A direct head-to-head between this automated pipeline and a panel of human analysts on the same AFO videos would clarify whether the system truly exceeds human consistency, and is a sensible next evaluation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents a computer-vision pipeline for estimating species composition of tuna catches on purse seiners from electronic monitoring (EM) video. It quantifies expert disagreement on bigeye/yellowfin identification, trains and compares segmentation models (Mask R-CNN, YOLOv9+SAM2, DINOv2+SAM2) and classification models (standard vs. hierarchical), and evaluates the full pipeline on 21 artificial fishing operations (AFOs) with known species composition. The authors report that YOLOv9+SAM2 achieves the best segmentation, and they claim that combining it with hierarchical classification yields the best composition estimates, with 84.8% of individuals segmented and a mean average error of 4.5%.

Significance. If the reported performance were supported by a fair, pre-specified comparison, the work would be a meaningful step toward automated EM-based monitoring in tropical tuna purse-seine fisheries, where manual video review is costly and species identification from EM images is notoriously difficult. Strengths include a purpose-built ground-truth dataset with onboard expert identification, repeated 5-fold cross-validation, and evaluation against fully known composition in AFOs. The expert-agreement experiment (Section 3.1) is a valuable result in itself, quantifying the difficulty of BET/YFT discrimination from EM imagery. However, the main headline claim is not sustained by the evidence as presented: the hierarchical classifier is not consistently superior, and the 4.5% figure results from post-hoc per-trip selection. The paper requires major revision to align its claims with the reported data.

major comments (4)
  1. [Section 3.4 / Table 3] The claim that hierarchical classification produced "the best estimations" with 4.5% MAE is not supported as stated. Table 3's caption states that "the best classification approach (i.e., standard or hierarchical) was selected for each trip," and the text reports that the standard classifier had lower overall MAE (10.5% vs. 17.8%) and lower trip-1 MAE (12.3% vs. 24.5%), with hierarchical better only on trip 2 (4.5% vs. 6.9%). Thus the headline 4.5% is an ex post minimum over two classifiers on one trip, not the performance of a pre-specified pipeline. The abstract's conclusion of "superior generalization" by the hierarchical approach is contradicted by Section 3.4 and should be revised.
  2. [Abstract / Section 3.4] The headline "84.8% of the individuals being segmented and classified with a mean average error of 4.5%" conflates two metrics from different tables and applies only to the second trip: 84.8% is the trip-2 segmentation proportion in Table 2 (YOLOv9+SAM2), and 4.5% is the trip-2 MAE in Table 3 (hierarchical classification). The paper does not report a joint measure of segmented-and-classified individuals; Table 3 MAEs are computed on segmented individuals only. Please report per-trip and per-classifier results separately and define a single, pre-specified evaluation protocol.
  3. [Table C.6 / Section 3.4] Segmentation rates above 100% (102.7% for AFO 2_01 and 105.2% for 2_06) indicate that the pipeline counts more individuals than the known total, implying double-counting or tracking artifacts. These cases are not discussed in the text, yet the MAE calculations rely on these counts. The authors should analyze and report counting errors (e.g., per-AFO counts after tracking) and correct or exclude double-counted segments, since this directly affects the reliability of the reported composition estimates.
  4. [Section 2.1 / Discussion] The test conditions are artificial: fish were removed from real catches, identified on deck, and returned to the conveyor belt in single-species or mixed batches. This likely alters stacking, occlusion, motion, and lighting relative to real seine hauls. The paper evaluates only 21 AFOs from two trips of one vessel, yet the Discussion claims "near-operational conditions." The external-validity risk should be explicitly acknowledged, and ideally the pipeline should be tested on at least a few unmodified real fishing operations to support the claim of accurate operational estimation.
minor comments (5)
  1. [Abstract] Typographical issues: "a integration" should be "an integration"; "YFT,Thunnus albacares" is missing a space after the comma.
  2. [Section 3.4 / Discussion] The Discussion states "Although our current seven validated operations demonstrate that the approach works," but Section 3.4 reports 21 AFOs from two trips (14 + 7). Clarify whether "seven" refers only to the second trip or whether some AFOs were excluded from the validation.
  3. [Section 2.3] The IoU threshold for a "valid" segmentation is not specified. The mAP calculation is described, but the recall and the reported "percentage of individuals segmented" depend on a threshold; specify the threshold and justify its choice.
  4. [Section 2.6 / Table A.4] The text says YOLOv9 was limited to five repetitions of cross-validation due to training time, while other models used 10 repetitions. This is fine, but the asymmetry should be stated in the main text and considered when comparing mAP/recall standard deviations.
  5. [Section 3.3 / Figure 5] In the hierarchical classification step 1, NO_TARGET accuracy is 0.69±0.11, much lower than in the standard classifier (0.90±0.03). The explanation given is plausible, but the text should also report how the NO_TARGET class is handled in the subsequent hierarchical steps, since the confusion matrices alone do not clarify this.

Circularity Check

1 steps flagged

The 'best estimations' claim is a post-hoc test-set minimum, not an independent prediction; the rest of the pipeline comparison is empirical and not circular.

specific steps
  1. fitted input called prediction [Abstract; Section 3.4 'Testing with artificial fishing operations', Table 3 caption]
    "Combining YOLOv9-SAM2 with the hierarchical classification produced the best estimations, with 84.8% of the individuals being segmented and classified with a mean average error of 4.5%. ... The best classification approach (i.e., standard or hierarchical) was selected for each trip."

    The headline 4.5% MAE is not the out-of-sample error of a pre-specified pipeline. The paper reports that the standard classifier had a lower overall MAE (10.5% vs 17.8%) and that the two classifiers were selected per trip after test results were known: standard won on trip 1 (12.3% vs 24.5%), hierarchical won on trip 2 (4.5% vs 6.9%). Thus 'hierarchical produced the best estimations' is true on the second trip by the selection rule itself, not by an independent test of superiority. Reporting the post-hoc minimum as the system's accuracy makes the claimed 'best' figure a selected test-set minimum rather than a prediction.

full rationale

This is an empirical computer-vision paper, not a closed-form derivation, and the main evaluation is against independent ground truth: fish were identified on board by experienced observers, models were trained on monospecific batches, and held-out mixed AFOs were used for testing. The 84.8% segmentation rate and the per-species MAE values are genuine measurements of model outputs versus that ground truth, so that part is not circular. The circular/selection issue is confined to the paper's central claim of 'hierarchical best': the caption of Table 3 admits the best classifier was chosen per trip after seeing test results, and the reported 4.5% MAE is the minimum of the two classifiers for the second trip. Calling this minimum a demonstration of hierarchical superiority reduces the claim to the selection criterion. Additional concerns—artificial fishing operations, 21 AFOs from one vessel, and segmentation rates above 100% in Table C.6—are data-quality and external-validity limitations, not circularity. On balance, the pipeline's measurements have independent content, but the headline 'best' figure is partly an artifact of post-hoc model selection, so the score is 4 rather than 0-2.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 0 invented entities

The pipeline is empirical and relies on the ground-truth protocol and AFO representativeness rather than on a closed-form derivation. No new physical entities are introduced. The main hidden choices are post-hoc classifier selection, the unspecified IoU threshold, and the expert-agreement cutoff.

free parameters (3)
  • Trip-wise classification model selection = per trip: standard for trip 1, hierarchical for trip 2
    Table 3 caption says 'The best classification approach ... was selected for each trip.' This post-hoc selection on test data is the basis for the abstract's 4.5% MAE claim.
  • IoU threshold for a valid segmentation
    Section 2.3 says a segment is valid when IoU exceeds 'a predefined threshold' but no value is given; reported segmentation percentages (Table 2) depend on this threshold.
  • Expert agreement inclusion threshold = at least 4 experts labeled the individual
    Section 3.1 computes inter-expert agreement only on individuals identified by at least four experts; changing this cutoff would change the reported 42.9%/57.1% values.
axioms (3)
  • domain assumption Onboard experts can reliably identify species by handling fish, even though image-based expert identification is unreliable.
    Section 2.1 relies on in-situ identification as ground truth; the expert image experiment shows this assumption is nontrivial and is not independently validated against genetics or morphometrics.
  • domain assumption Artificial fishing operations (AFOs) are representative of real fishing operations for evaluating catch composition.
    Section 2.1/3.4: fish were removed, identified, and placed back on the conveyor belt in controlled batches, changing occlusion, stacking, and handling compared with real catch flow.
  • domain assumption ByteTrack with optical-flow gating counts each individual fish exactly once.
    Section 2.4 describes tracking; Table C.6 shows segmentation rates above 100% in several AFOs, contradicting this assumption.

pith-pipeline@v1.3.0-alltime-deepseek · 26604 in / 10895 out tokens · 97167 ms · 2026-08-04T06:44:45.915221+00:00 · methodology

0 comments
read the original abstract

Purse seiners play a crucial role in tuna fishing, as approximately 69% of the world's tropical tuna is caught using this gear. All tuna Regional Fisheries Management Organizations have established minimum standards to use electronic monitoring (EM) in fisheries in addition to traditional observers. The EM systems produce a massive amount of video data that human analysts must process. Integrating artificial intelligence (AI) into their workflow can decrease that workload and improve the accuracy of the reports. However, species identification still poses significant challenges for AI, as achieving balanced performance across all species requires appropriate training data. Here, we quantify the difficulty experts face to distinguish bigeye tuna (BET, Thunnus Obesus) from yellowfin tuna (YFT, Thunnus Albacares) using images captured by EM systems. We found inter-expert agreements of 42.9% $\pm$ 35.6% for BET and 57.1% $\pm$ 35.6% for YFT. We then present a multi-stage pipeline to estimate the species composition of the catches using a reliable ground-truth dataset based on identifications made by observers on board. Three segmentation approaches are compared: Mask R-CNN, a combination of DINOv2 with SAM2, and a integration of YOLOv9 with SAM2. We found that the latest performs the best, with a validation mean average precision of 0.66 $\pm$ 0.03 and a recall of 0.88 $\pm$ 0.03. Segmented individuals are tracked using ByteTrack. For classification, we evaluate a standard multiclass classification model and a hierarchical approach, finding a superior generalization by the hierarchical. All our models were cross-validated during training and tested on fishing operations with fully known catch composition. Combining YOLOv9-SAM2 with the hierarchical classification produced the best estimations, with 84.8% of the individuals being segmented and classified with a mean average error of 4.5%.

Figures

Figures reproduced from arXiv: 2511.15468 by Ahmad Kamal, Ignacio Arganda-Carreras, I\~naki Quincoces, Izaro Goienetxea, Jaime Valls Miro, Jon Ruiz, Jose A. Fernandes-Salvador, Xabier Lekunberri.

Figure 1
Figure 1. Figure 1: Workflow that outlines the pipeline for fish image analysis. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Manually annotated image. This image contains a masks for each fish on the conveyor belt that is visible from the point of view of the camera. This frame belongs to a monospecific SKJ sample. It was used during the training of the models. 2.3 Automatic image segmentation. The con￾ventional approach to segmentation and classifi￾cation addresses both tasks in a single stage, as￾signing distinct labels to eac… view at source ↗
Figure 3
Figure 3. Figure 3: Different statistics for the analysis based on expert identifi [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Comparison of all the segmentation approaches. [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Confusion matrices for the four classification models. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

64 extracted references · 18 canonical work pages

  1. [1]

    The State of World Fisheries and Aquaculture 2024 – Blue Transformation in action

    FAO. The State of World Fisheries and Aquaculture 2024 – Blue Transformation in action. June 2024. doi:10.4060/cd0683en

  2. [2]

    Snapshot of the Large-Scale Tropical Tuna Purse Seine Fish- ing Fleets as of June 2024.International Seafood Sustainability Foundation, June 2024

    Ana Justel-Rubio. Snapshot of the Large-Scale Tropical Tuna Purse Seine Fish- ing Fleets as of June 2024.International Seafood Sustainability Foundation, June 2024

  3. [3]

    Griffiths, Valerie Allain, Simon D

    Shane P . Griffiths, Valerie Allain, Simon D. Hoyle, Tim A. Lawson, and Simon J. Nicol. Just a FAD? Ecosystem impacts of tuna purse-seine fishing associated with fish aggregating devices in the western Pacific Warm Pool Province.Fish- eries Oceanography, 28(1):94–112, January 2019. ISSN 1054-6006, 1365-2419. doi:10.1111/fog.12389

  4. [4]

    Duparc, Pascal Cauquil, Mathieu Depetris, Patrice Dewals, Daniel Gaertner, A

    A. Duparc, Pascal Cauquil, Mathieu Depetris, Patrice Dewals, Daniel Gaertner, A. Hervé, Julien Lebranchu, Francis Marsac, P . Bach, and Pascal Bach. Assess- ment of accuracy in processing purse seine tropical tuna catches with the T3 methodology using French fleet data. page 1, 2018

  5. [5]

    Antoine Duparc, V. Aragno, Mathieu Depetris, Laurent Floch, Pascal Cauquil, Julien Lebranchu, Daniel Gaertner, Francis Marsac, Pascal Bach, Mathieu De- petris, Laurent Floch, Pascal Cauquil, Julien Lebranchu, Daniel Gaertner, Francis 12 X. Lekunberriet al.| Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners 3.4 Testin...

  6. [6]

    At-Sea Observing Using Video-Based Electronic Monitoring

    Howard McElderry. At-Sea Observing Using Video-Based Electronic Monitoring. January 2008

  7. [7]

    Designing and Implementing Electronic Monitoring Systems for Fisheries: A Sup- plement to the Catch Share Design Manual

    Rod Fujita, Christopher Cusack, Rachel Karasik, and Helen Takade-Heumacher. Designing and Implementing Electronic Monitoring Systems for Fisheries: A Sup- plement to the Catch Share Design Manual. Technical report, Environmental Defense Fund, San Francisco, 2018

  8. [8]

    Aloysius T. M. Van Helmond, Lars O. Mortensen, Kristian S. Plet-Hansen, Clara Ulrich, Coby L. Needle, Daniel Oesterwind, Lotte Kindt-Larsen, Thomas Catch- pole, Stephen Mangi, Christopher Zimmermann, Hans Jakob Olesen, Nick Bai- ley, Heidrikur Bergsson, Jørgen Dalskov, Jon Elson, Malo Hosken, Lisa Peterson, Howard McElderry, Jon Ruiz, Johanna P . Pierre, ...

  9. [9]

    Comparing electronic monitoring system with ob- server data for estimating non-target species and discards on french tropical tuna purse seine vessels

    K Briand, A Bonnieux, W Le Dantec, S Le Couls, P Bach, A Maufroy, A Relot- Stirnemann, and P Sabarros. Comparing electronic monitoring system with ob- server data for estimating non-target species and discards on french tropical tuna purse seine vessels. (Col Vol Sci Pap ICCAT 74):3813–3831, 2018

  10. [10]

    Sabarros, Gwenaëlle Wain, An- toine Bonnieux, Sarah Le Couls, Romain Godefroy, Patrick Moëlo, Tiphaine Bet- tali, Michel Goujon, and Julien Lebranchu

    Karine Briand, Alexandra Maufroy, Philippe S. Sabarros, Gwenaëlle Wain, An- toine Bonnieux, Sarah Le Couls, Romain Godefroy, Patrick Moëlo, Tiphaine Bet- tali, Michel Goujon, and Julien Lebranchu. The feasibility and challenges of col- lecting Electronic Monitoring System (EMS) data on French and associated purse seiners in relation to IOTC minimum standa...

  11. [11]

    Increasing the functionalities and accuracy of fisheries electronic monitoring systems.Aquatic Conservation: Marine and Freshwater Ecosystems, 29(6):901–926, 2019

    Eric Gilman, Gonzalo Legorburu, Andrew Fedoruk, Craig Heberer, Mark Zim- ring, and Amos Barkai. Increasing the functionalities and accuracy of fisheries electronic monitoring systems.Aquatic Conservation: Marine and Freshwater Ecosystems, 29(6):901–926, 2019. ISSN 1099-0755. doi:10.1002/aqc.3086

  12. [12]

    Preliminary study about the suitability of an electronic monitoring system to record scientific and other information from the tropical tuna purse seine fishery

    J P Monteagudo, G Legorburu, A Justel-Rubio, and V Restrepo. Preliminary study about the suitability of an electronic monitoring system to record scientific and other information from the tropical tuna purse seine fishery. page 20, 2014

  13. [13]

    Murua, M

    H. Murua, M. Herrera, J. Morón, F . Abascal, G. Legoburu, M. Roman, G. Moreno, M. Hosken, and V. Restrepo. Comparing Electronic Monitoring and human ob- server collected fishery data in the tropical tuna purse seine operating in the Pacific Ocean.IATTC SAC-11 INF-G, 2020

  14. [14]

    Minimum Stan- dards for Electronic Monitoring Systems in Tropical Tuna Purse Seine and Long- line Fisheries

    Hilario Murua, Jon Ruiz, Ana Justel-Rubio, and Victor Restrepo. Minimum Stan- dards for Electronic Monitoring Systems in Tropical Tuna Purse Seine and Long- line Fisheries. Technical report, International Seafood Sustainability Foundation, Washington, D.C., USA, September 2022

  15. [15]

    J. Ruiz, A. Batty, P . Chavance, H. McElderry, V. Restrepo, P . Sharples, J. Santos, and A. Urtizberea. Electronic monitoring trials on in the tropical tuna purse-seine fishery.ICES Journal of Marine Science, 72(4):1201–1213, May 2015. ISSN 1054-3139. doi:10.1093/icesjms/fsu224

  16. [16]

    J. Ruiz, I. Krug, O. Gonzalez, and G. Hammann. E-EYE PLUS: Electronic moni- toring trial on the tropical tuna purse seine fleet.ICCAT, Madrid, Spain, 3, 2016

  17. [17]

    Comments on the assessment of catch by species in the tropical purse seine fishery

    Jon Ruiz, Antoine Duparc, Francisco Abascal, Pascal Bach, Jose Carlos Baez, Pedro Pascual, Daniel Gaertner, Francis Marsac, and Josu Santiago. Comments on the assessment of catch by species in the tropical purse seine fishery. 2021

  18. [18]

    Fast freezing, slow tempering: The key to tuna (Thunnus albacares) quality retention.International Journal of Food Science and Technology, 59(5):3031– 3044, May 2024

    Zhenze Duanmu, Y ajin Zhang, Feng Li, Y ao Y ao, Hu Shi, Fanbin Kong, and Y ang Jiao. Fast freezing, slow tempering: The key to tuna (Thunnus albacares) quality retention.International Journal of Food Science and Technology, 59(5):3031– 3044, May 2024. ISSN 0950-5423. doi:10.1111/ijfs.17034

  19. [19]

    Miquel Palmer, Amaya Álvarez Ellacuría, Vicenç Moltó, and Ignacio A. Catalán. Automatic, operational, high-resolution monitoring of fish length and catch num- bers from landings using deep learning.Fisheries Research, 246:106166, Febru- ary 2022. ISSN 0165-7836. doi:10.1016/j.fishres.2021.106166

  20. [20]

    Mask R-CNN

    Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask R-CNN. arXiv:1703.06870 [cs.CV], January 2018. doi:10.48550/arxiv.1703.06870

  21. [21]

    Maria Sokolova, Manuel Cordova, Henk Nap, Aloysius van Helmond, Michiel Mans, Arjan Vroegop, Angelo Mencarelli, and Gert Kootstra. An integrated end- to-end deep neural network for automated detection of discarded fish species and their weight estimation.ICES Journal of Marine Science, 80(7):1911–1922, September 2023. ISSN 1054-3139. doi:10.1093/icesjms/fsad118

  22. [22]

    Multi-stage image-based approach for fish detection and weight estimation.Biosystems Engineering, 257:104239, September 2025

    Manuel Córdova, Maria Sokolova, Aloysius van Helmond, Angelo Mencarelli, and Gert Kootstra. Multi-stage image-based approach for fish detection and weight estimation.Biosystems Engineering, 257:104239, September 2025. ISSN 1537-

  23. [23]

    Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S

    Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Ni- ladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Dem- szky, Chris Donahue, Moussa Doumbouya, Esin Durmus, ...

  24. [24]

    SAM2 for Image and Video Segmentation: A Comprehensive Survey, March 2025

    Zhang Jiaxing and Tang Hao. SAM2 for Image and Video Segmentation: A Comprehensive Survey, March 2025. arXiv:2503.12781 [cs]

  25. [25]

    DINOv2: Learning Robust Visual Features without Super- vision, February 2024

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El- Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po- Y ao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick ...

  26. [26]

    SAM 2: Segment Anything in Images and Videos, October 2024

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, and Christoph Feichtenhofer. SAM 2: Segment Anything in Images and Videos, October 2024. arXiv:2408.00714 [cs]

  27. [27]

    Training- free fish species population monitoring in unconstrained underwater videos

    Isaak Kavasidis, Amelia Sorrenti, Orazio Tomarchio, Daniela Giordano, Marco Milazzo, Gabriele Turco, Carlo Cattano, and Concetto Spampinato. Training- free fish species population monitoring in unconstrained underwater videos. Procedia Computer Science, 257:809–816, January 2025. ISSN 1877-0509. doi:10.1016/j.procs.2025.03.104

  28. [28]

    Evaluation of Segment Anything Model 2: The Role of SAM2 in the Underwater Environment, August 2024

    Shijie Lian and Hua Li. Evaluation of Segment Anything Model 2: The Role of SAM2 in the Underwater Environment, August 2024. arXiv:2408.02924 [cs]

  29. [29]

    Advancing Fisheries Re- search and Management with Computer Vision: A Survey of Recent Develop- ments and Pending Challenges.Fishes, 10(2):74, February 2025

    Jesse Eickholt, Jonathan Gregory, and Kavya Vemuri. Advancing Fisheries Re- search and Management with Computer Vision: A Survey of Recent Develop- ments and Pending Challenges.Fishes, 10(2):74, February 2025. ISSN 2410-

  30. [30]

    Fernandes

    Xabier Lekunberri, Jon Ruiz, Iñaki Quincoces, Fadi Dornaika, Ignacio Arganda- Carreras, and Jose A. Fernandes. Identification and measurement of trop- ical tuna species in purse seiner catches using computer vision and deep learning.Ecological Informatics, 67:101495, March 2022. ISSN 15749541. doi:10.1016/j.ecoinf.2021.101495

  31. [31]

    ai Corporation

    CVAT. ai Corporation. Computer Vision Annotation Tool (CVAT), July 2024

  32. [32]

    Bett, Andrew R

    Tammy Horton, Leigh Marsh, Brian J. Bett, Andrew R. Gates, Daniel O. B. Jones, Noëlie M. A. Benoist, Simone Pfeifer, Erik Simon-Lledó, Jennifer M. Durden, Leen Vandepitte, and Ward Appeltans. Recommendations for the Standardisation of Open Taxonomic Nomenclature for Image-Based Identifi- cations.Frontiers in Marine Science, 8, February 2021. ISSN 2296-774...

  33. [33]

    ASFIS List of Species for Fishery Statistics Purposes

    FAO. ASFIS List of Species for Fishery Statistics Purposes. In: Fisheries and Aquaculture., 2024

  34. [34]

    YOLOv9: Learning What Y ou Want to Learn Using Programmable Gradient Information, February

    Chien-Y ao Wang, I.-Hau Y eh, and Hong-Yuan Mark Liao. YOLOv9: Learning What Y ou Want to Learn Using Programmable Gradient Information, February

  35. [35]

    DINOSim: Zero-Shot Object Detection and Semantic Segmentation on Electron Microscopy Images, March 2025

    Aitor González-Marfil, Estibaliz Gómez-de Mariscal, and Ignacio Arganda- Carreras. DINOSim: Zero-Shot Object Detection and Semantic Segmentation on Electron Microscopy Images, March 2025. Pages: 2025.03.09.642092 Sec- tion: New Results

  36. [36]

    Ni, and Heung-Y eung Shum

    Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M. Ni, and Heung-Y eung Shum. DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection, July 2022. arXiv:2203.03605 [cs]

  37. [37]

    Faster R-CNN: To- wards Real-Time Object Detection with Region Proposal Networks.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 39(6):1137–1149, June

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: To- wards Real-Time Object Detection with Region Proposal Networks.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 39(6):1137–1149, June

  38. [38]

    Simple Online and Realtime Tracking, July 2017

    Alex Bewley, Zongyuan Ge, Lionel Ott, Fabio Ramos, and Ben Upcroft. Simple Online and Realtime Tracking, July 2017. arXiv:1602.00763

  39. [39]

    Simple Online and Realtime Tracking with a Deep Association Metric, March 2017

    Nicolai Wojke, Alex Bewley, and Dietrich Paulus. Simple Online and Realtime Tracking with a Deep Association Metric, March 2017. arXiv:1703.07402

  40. [40]

    ByteTrack: Multi-object Tracking by Associating Every Detection Box

    Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo, Wenyu Liu, and Xinggang Wang. ByteTrack: Multi-object Tracking by Associating Every Detection Box. In Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner, editors,Computer Vision – ECCV 2022, pages 1–21, Cham, 2022. Springer Nature Switz...

  41. [41]

    Mahmoud Abouelyazid. Comparative Evaluation of SORT, DeepSORT, and Byte- Track for Multiple Object Tracking in Highway Videos.International Journal of Sustainable Infrastructure for Cities and Societies, 8(11):42–52, November 2023. Number: 11

  42. [42]

    Jumabek Alikhanov and Hakil Kim. Online Action Detection in Surveillance Sce- narios: A Comprehensive Review and Comparative Study of State-of-the-Art Multi-Object Tracking Methods.IEEE Access, 11:68079–68092, 2023. ISSN 2169-3536. doi:10.1109/ACCESS.2023.3292539. Conference Name: IEEE Ac- cess

  43. [43]

    Dynamic Target Tracking and Fol- lowing with UAVs Using Multi-Target Information: Leveraging YOLOv8 and MOT Algorithms.DRONES, 8(9):488, September 2024

    Diogo Ferreira and Meysam Basiri. Dynamic Target Tracking and Fol- lowing with UAVs Using Multi-Target Information: Leveraging YOLOv8 and MOT Algorithms.DRONES, 8(9):488, September 2024. ISSN 2504-446X. doi:10.3390/drones8090488. Num Pages: 23 Place: Basel Publisher: MDPI Web of Science ID: WOS:001324028000001

  44. [44]

    Designing Network Design Spaces, March 2020

    Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár. Designing Network Design Spaces, March 2020. arXiv:2003.13678 [cs]

  45. [45]

    Mujtaba and Nihar R

    Dena F . Mujtaba and Nihar R. Mahapatra. Hierarchical Deep Learning Models X. Lekunberriet al.| Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners 13 for Identification of Fish Species.2022 International Conference on Computa- tional Science and Computational Intelligence (CSCI), pages 1558–1562, De- cember 2022. doi:...

  46. [46]

    Ackland, David M

    Sarah J. Ackland, David M. Richardson, and Tamara B. Robinson. A Method for Conveying Confidence in iNaturalist Observations: A Case Study Using Non-Native Marine Species.Ecology and Evolution, 14(10): e70376, 2024. ISSN 2045-7758. doi:10.1002/ece3.70376. _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/ece3.70376

  47. [47]

    Shepherd, Juliet A

    Thomas Mesaglio, Kelly A. Shepherd, Juliet A. Wege, Russell L. Barrett, Hervé Sauquet, and Will K. Cornwell. Expert identification blitz: A rapid high value approach for assessing and improving iNaturalist identification ac- curacy and data precision and confidence.PLANTS, PEOPLE, PLANET, 7 (5):1469–1484, 2025. ISSN 2572-2611. doi:10.1002/ppp3.70005. _epr...

  48. [48]

    Culverhouse, Robert Williams, Beatriz Reguera, Vincent Herry, and Son- soles González-Gil

    Phil F . Culverhouse, Robert Williams, Beatriz Reguera, Vincent Herry, and Son- soles González-Gil. Do experts make mistakes? A comparison of human and machine identification of dinoflagellates.Marine Ecology Progress Series, 247: 17–25, February 2003. ISSN 0171-8630, 1616-1599. doi:10.3354/meps247017

  49. [49]

    Culverhouse

    Phil F . Culverhouse. Human and machine factors in algae monitoring perfor- mance.Ecological Informatics, 2(4):361–366, December 2007. ISSN 1574-9541. doi:10.1016/j.ecoinf.2007.07.001

  50. [50]

    Austen, Markus Bindemann, Richard A

    Gail E. Austen, Markus Bindemann, Richard A. Griffiths, and David L. Roberts. Species identification by conservation practitioners using online images: accu- racy and agreement between experts.PeerJ, 6:e4157, January 2018. ISSN 2167-8359. doi:10.7717/peerj.4157. Publisher: PeerJ Inc

  51. [51]

    Foundation Models

    Johannes Schneider, Christian Meske, and Pauline Kuss. Foundation Models. Business & Information Systems Engineering, 66(2):221–231, April 2024. ISSN 1867-0202. doi:10.1007/s12599-024-00851-0

  52. [52]

    The Fishnet Open Images Database: A Dataset for Fish Detection and Fine-Grained Categorization in Fisheries, June 2021

    Justin Kay and Matt Merrifield. The Fishnet Open Images Database: A Dataset for Fish Detection and Fine-Grained Categorization in Fisheries, June 2021. arXiv:2106.09178 [cs]

  53. [53]

    Rick van Essen, Angelo Mencarelli, Aloysius van Helmond, Linh Nguyen, Ju- rgen Batsleer, Jan-Jaap Poos, and Gert Kootstra. Automatic discard reg- istration in cluttered environments using deep learning and object tracking: class imbalance, occlusion, and a comparison to human review.ICES Jour- nal of Marine Science, 78(10):3834–3846, December 2021. ISSN 1...

  54. [54]

    RSNC-YOLO: A Deep-Learning-Based Method for Automatic Fine-Grained Tuna Recognition in Complex Environments.Applied Sciences, 14 (22):10732, January 2024

    Wenjie Xu, Hui Fang, Shengchi Yu, Shenglong Y ang, Haodong Y ang, Yujia Xie, and Y ang Dai. RSNC-YOLO: A Deep-Learning-Based Method for Automatic Fine-Grained Tuna Recognition in Complex Environments.Applied Sciences, 14 (22):10732, January 2024. ISSN 2076-3417. doi:10.3390/app142210732. Pub- lisher: Multidisciplinary Digital Publishing Institute

  55. [55]

    Live Weight Prediction of Cattle Based on Deep Re- gression of RGB-D Images.Agriculture, 12(11):1794, November 2022

    Alexey Ruchay, Vitaly Kober, Konstantin Dorofeev, Vladimir Kolpakov, Alexey Gladkov, and Hao Guo. Live Weight Prediction of Cattle Based on Deep Re- gression of RGB-D Images.Agriculture, 12(11):1794, November 2022. ISSN 2077-0472. doi:10.3390/agriculture12111794. Publisher: Multidisciplinary Digi- tal Publishing Institute

  56. [56]

    A brief survey on RGB-D semantic segmentation using deep learning.Displays, 70:102080, December 2021

    Changshuo Wang, Chen Wang, Weijun Li, and Haining Wang. A brief survey on RGB-D semantic segmentation using deep learning.Displays, 70:102080, December 2021. ISSN 0141-9382. doi:10.1016/j.displa.2021.102080

  57. [57]

    International Commission for the Conservation of Atlantic Tuna, 2006

    ICCAT.ICCAT Manual. International Commission for the Conservation of Atlantic Tuna, 2006. ISBN 978-92-990055-0-7

  58. [58]

    Fernandes-Salvador, Nicolas Goñi, Igor Granado, Iñaki Quincoces, Leire Ibaibarriaga, Jon Ruiz, Hilario Murua, and Ainhoa Caballero

    Nerea Goikoetxea, Izaro Goienetxea, Jose A. Fernandes-Salvador, Nicolas Goñi, Igor Granado, Iñaki Quincoces, Leire Ibaibarriaga, Jon Ruiz, Hilario Murua, and Ainhoa Caballero. Machine-learning aiding sustainable Indian Ocean tuna purse seine fishery.Ecological Informatics, 81:102577, July 2024. ISSN 1574-9541. doi:10.1016/j.ecoinf.2024.102577

  59. [59]

    Global habitat preferences of commercially valuable tuna.Deep Sea Research Part II: Topical Studies in Oceanography, 113:102–112, March 2015

    Haritz Arrizabalaga, Florence Dufour, Laurence Kell, Gorka Merino, Leire Ibaibar- riaga, Guillem Chust, Xabier Irigoien, Josu Santiago, Hilario Murua, Igaratza Fraile, Marina Chifflet, Nerea Goikoetxea, Y olanda Sagarminaga, Olivier Aumont, Laurent Bopp, Miguel Herrera, Jean Marc Fromentin, and Sylvain Bonhomeau. Global habitat preferences of commercially...

  60. [645]

    doi:10.1016/j.dsr2.2014.07.001. 14 X. Lekunberriet al.| Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners 3.4 Testing with artificial fishing operations Supplementary Material A. Reproducibility Table A.4.Summary of computational resources, data augmentation parameters, hyperparameter configurations, and training dur...

  61. [2017]

    doi:10.1109/TPAMI.2016.2577031

    ISSN 0162-8828, 2160-9292. doi:10.1109/TPAMI.2016.2577031

  62. [2024]

    arXiv:2402.13616 [cs]

  63. [3888]

    Number: 2 Publisher: Multidisciplinary Digi- tal Publishing Institute

    doi:10.3390/fishes10020074. Number: 2 Publisher: Multidisciplinary Digi- tal Publishing Institute

  64. [5110]

    doi:10.1016/j.biosystemseng.2025.104239