REVIEW 3 major objections 6 minor 69 references
Improving Generalization Performance of YOLOv8 for Camera Trap Object Detection
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This thesis claims that adding a Global Attention Mechanism, a modified multi-scale feature fusion, and WIoUv3 loss to YOLOv8s improves generalization to never-before-seen camera trap locations, with Trans-Test mAP50 rising from 0.520 to…
desk verdict The claimed trans-location gain is a narrow, plausible result that needs seeds and a proper fusion ablation before anyone should lean on it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is three coordinated modifications to YOLOv8s. The Global Attention Mechanism (GAM) is an attention block with a channel submodule (3D permutation plus a two-layer MLP) and a spatial submodule (two 7x7 convolutions); inserted at layer 9 it reweights features so the detector emphasizes object properties and suppresses background. The modified feature fusion adds the early C2f block's output at layer 2 to the neck's upsampling and concatenation path, preserving small-object detail. WIoUv3 replaces CIoU as the bounding-box regression loss; its dynamic non-monotonic focusing mechanism assigns small gradient gains both to very low- and very high-quality boxes, letting ordinary boxes drive optimization. The paper's argument is that each modification targets one of the three failure modes it identifies in the baseline: background leakage, lost fine-grained localization detail, and gradient suppression by numerous overlaps.
What would settle it
Train the eight combinations of the three modifications on the same camera trap benchmark with multiple seeds and with the empty category kept in the test sets, then compare Trans-Test mAP50; the central claim fails if the full model does not consistently beat the baseline.
Extended reading notes
Core claim
The paper's central claim is that a bundle of three architectural and loss changes improves generalization of YOLOv8s to novel camera trap locations. GAM, placed after the backbone's layer 9, suppresses background activation; feeding the C2f output of layer 2 into the neck preserves fine spatial detail lost by downsampling; and WIoUv3 down-weights low-quality bounding boxes during regression. The reported evidence is the Trans-Test mAP50 rising from 0.520 in the baseline to 0.541 in the improved model, together with Grad-CAM heatmaps showing the improved model's activations concentrated on the animal rather than the background. The paper also reports the trade: same-location Cis-Test mAP50 falls from 0.813 to 0.772.
Load-bearing premise
The central claim assumes that the Trans-Test mAP50 difference is a real measure of generalization and that the package of three modifications, rather than training noise or the removal of empty images, caused it.
Editorial extensions
If this is right
- If the result holds, a detector trained on camera traps from one set of locations can be deployed at new sites with roughly a 4% relative gain in mAP50 over the stock model, without any training at the new site.
- The ablation indicates the components interact: WIoUv3 alone improves Trans-Test mAP50 (0.528), GAM alone lowers it (0.496), and the combination reaches 0.541, so the attention module's benefit appears contingent on the loss change.
- The improved model trades some same-location accuracy (Cis-Test mAP50 drops from 0.813 to 0.772) for better cross-location transfer, which is the intended trade when deployment sites are unknown.
- The heatmap evidence suggests the improvement comes with a visible shift in what the network attends to: background activation seen in the baseline is suppressed in the improved model.
Reading between the lines
- The paper does not isolate the modified multi-scale feature fusion in its ablation, so the individual contribution of adding the layer-2 C2f features to the neck is untested; a factorial ablation toggling each component would be needed to attribute the gain.
- Because all empty images were removed from every set, the evaluation does not measure the false-trigger suppression that motivates the work; re-adding the empty category could change the precision and mAP numbers.
- Each configuration appears to have been trained once; with typical YOLO run-to-run variance at the scale of the observed 0.021 mAP50 gap, seed-averaged runs are needed to confirm the claimed gain is not noise.
- The heatmaps compare different layers (layer 21 in the baseline versus layer 28 in the improved model), so the visual claim that attention shifted from background to object would be stronger if the same layer were compared.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The thesis-style paper proposes three modifications to YOLOv8s for camera-trap object detection: inserting a Global Attention Mechanism (GAM) after the backbone, adding the C2f layer-2 feature map to the neck for modified multi-scale feature fusion, and replacing CIoU with the WIoUv3 bounding-box regression loss. Using a subset of the Caltech Camera Traps dataset with a held-out trans-location test set, the paper reports that the improved model reaches Trans-Test mAP50 of 0.541 versus 0.520 for the baseline, while Cis-Test mAP50 drops from 0.813 to 0.772. The contribution is presented as improved generalization to novel locations, supported by ablation experiments, Grad-CAM heatmaps, and qualitative examples on an internet-collected custom dataset.
Significance. If the reported improvement is reproducible across random seeds, the paper would provide a modest but useful demonstration that attention and loss-function changes can improve cross-location generalization on a standard camera-trap benchmark. The experimental design has genuine strengths: the trans-location test set is genuinely held out; the WIoUv3 hyperparameters are adopted from the original publication rather than tuned on the test set; and code and processed data are publicly linked. However, the central empirical claim is currently supported by one training run per configuration and by an ablation that never isolates one of the three proposed modifications, so the quantitative conclusion should be treated as preliminary until the missing evidence is supplied.
major comments (3)
- [§6.3, §6.4, §6.8 (Tables 6.6 and 6.7)] The headline generalization gain is based on a single training run per configuration. No random seed, repeated run, confidence interval, or significance test is reported anywhere in Chapter 6. The Trans-Test mAP50 difference is 0.021 (0.520 vs 0.541), while the same change reduces Cis-Test mAP50 by 0.041 (0.813 vs 0.772); the observed Trans-Test mAP50-95 gain is only 0.014. Given 180 epochs of stochastic training with augmentation, a 0.021 mAP50 gap is plausibly within run-to-run variance, and the drop on the in-domain split is larger than the claimed transfer gain. Please report results over at least three to five seeds, with mean and standard deviation or per-seed values, and a paired comparison or equivalent test. Without this, the Section 6.8 conclusion that the improved model performs much better on Trans-Test is not supported.
- [§4.2, §6.7, §6.8 (Table 6.8)] The ablation study does not isolate the modified multi-scale feature fusion. The rows are baseline, +WIoUv3, +GAM, and +GAM+WIoUv3; none varies the fusion component, and the full model evaluated in §6.4 includes all three modifications. If the last row is intended to include the modified fusion, the label does not say so, and the §6.8 discussion attributes the Trans-Test result to GAM and WIoUv3 only. As a result, the paper cannot attribute the reported gain to the package of three enhancements. Please add ablations that turn the fusion component on and off while keeping the other components fixed, and make the row labels explicit.
- [§6.1 (Tables 6.1 and 6.2), §6.8] All images without bounding-box annotations are removed from every split, including the test sets. Because empty frames are one of the defining challenges of camera-trap data (Section 1.2.1), the reported Trans-Test mAP measures a filtered object-detection task rather than end-to-end generalization to real camera-trap deployments. This does not necessarily invalidate the relative comparison, but it materially limits the real-world generalization claim. Please report an evaluation that includes the empty category, or explicitly justify and discuss the filtering as a limitation.
minor comments (6)
- [§6.8] The text states that the improved model achieves an improvement of 4%; since mAP50 rises from 0.520 to 0.541, the relative improvement is 4.0% while the absolute improvement is 0.021 mAP points. Please state both explicitly to avoid ambiguity.
- [§1.9] The thesis outline says 'You Only Live Once (YOLO) v8'; this should read 'You Only Look Once.'
- [§6.6 (Figures 6.19–6.24)] The heatmap comparison is qualitative and compares different layers (layer 21 for the baseline, layer 28 for the improved model) on only three images; please describe it as illustrative, or add quantitative measures of background activation.
- [§4.2] The modified feature fusion is described only as incorporating the output of the C2f module at Layer 2 into the neck; please specify the exact connection points, channel alignment, and resulting changes to the FPN/PAN structure so the architecture is reproducible without consulting the repository.
- [§6.9] The custom internet dataset evaluation is anecdotal; please report the number of images and quantitative metrics, or label it explicitly as a qualitative sanity check.
- [§5.2.6, Eq. (5.3)] The AP formula is not clear as typeset; please write the standard 101-point interpolation in conventional notation.
Circularity Check
No circularity found: the empirical claim is benchmarked against an external held-out Trans-Test set, and the added components are cited from independent prior work rather than derived from the paper's own outputs.
full rationale
The paper compares a baseline YOLOv8s with an improved variant (GAM attention, layer-2 feature fusion, WIoUv3 loss) on the Caltech Camera Traps subset of Beery et al., with a Trans-Test split from novel locations. The headline result, mAP50 rising from 0.520 to 0.541 on Trans-Test, is computed on an externally held-out test set that was not used to fit any parameter in the paper. The WIoUv3 hyperparameters (alpha=1.9, delta=3) are explicitly taken from the WIoU paper [56], and the GAM module is taken from [31]; neither is tuned on the test set and neither is a prior result of the present author. The modified feature fusion is described as adding the layer-2 C2f output into the neck, an architectural change rather than a redefinition of the metric. No load-bearing self-citation appears: the author's own contributions are the code repository and dataset split links in Appendix A, which are not used to justify the empirical result. There are legitimate methodological concerns, such as single training runs per configuration, no error bars, and the fact that the final model was selected after inspecting Trans-Test performance, but those are statistical validity issues rather than circular reasoning. The paper does not define its predictions in terms of its inputs, fit a parameter to a subset and then rename it as a prediction, or invoke any author-imported uniqueness theorem. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- WIoUv3 alpha =
1.9
- WIoUv3 delta =
3
- GAM spatial reduction ratio r =
unspecified
- Evaluation confidence and IoU thresholds =
confidence 0.25, IoU 0.45
assumptions (4)
- domain assumption The trans-location test split is a valid operational proxy for real-world generalization.
- ad hoc to paper Adding the C2f layer-2 output to the neck improves multi-scale fusion without introducing harmful noise.
- domain assumption Grad-CAM heatmaps faithfully indicate the features used for detection.
- domain assumption One 180-epoch training run per configuration is representative of that configuration's performance.
Cite this review
Pith. "Pith review of Improving Generalization Performance of YOLOv8 for Camera Trap Object Detection." pith.science (2026). https://pith.science/paper/UXS7ZLPW
@misc{pith2026241214211,
author = {Pith},
title = {Pith review of: Improving Generalization Performance of YOLOv8 for Camera Trap Object Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/UXS7ZLPW}},
note = {Machine review of arXiv:2412.14211}
}
read the original abstract
Camera traps have become integral tools in wildlife conservation, providing non-intrusive means to monitor and study wildlife in their natural habitats. The utilization of object detection algorithms to automate species identification from Camera Trap images is of huge importance for research and conservation purposes. However, the generalization issue, where the trained model is unable to apply its learnings to a never-before-seen dataset, is prevalent. This thesis explores the enhancements made to the YOLOv8 object detection algorithm to address the problem of generalization. The study delves into the limitations of the baseline YOLOv8 model, emphasizing its struggles with generalization in real-world environments. To overcome these limitations, enhancements are proposed, including the incorporation of a Global Attention Mechanism (GAM) module, modified multi-scale feature fusion, and Wise Intersection over Union (WIoUv3) as a bounding box regression loss function. A thorough evaluation and ablation experiments reveal the improved model's ability to suppress the background noise, focus on object properties, and exhibit robust generalization in novel environments. The proposed enhancements not only address the challenges inherent in camera trap datasets but also pave the way for broader applicability in real-world conservation scenarios, ultimately aiding in the effective management of wildlife populations and habitats.
Figures
Figures from the paper (65 more)
Reference graph
Works this paper leans on
-
[1]
Recognition in terra incognita
Sara Beery, Grant Van Horn, and Pietro Perona. Recognition in terra incognita. In Proceedings of the European conference on computer vision (ECCV) , pages 456–473, 2018
work page 2018
-
[2]
Efficient pipeline for camera trap image review
Sara Beery, Dan Morris, and Siyu Yang. Efficient pipeline for camera trap image review. arXiv preprint arXiv:1907.06772 , 2019
arXiv 1907
-
[3]
Marc Besson, Jamie Alison, Kim Bjerge, Thomas E. Gorochowski, Toke T. Høye, Tommaso Jucker, Hjalte M. R. Mann, and Christopher F. Clements. Towards the fully automated monitoring of ecological communities. Ecology Letters, 25(12):2753–2775,
-
[4]
J. David Blount, Mark W. Chynoweth, Austin M. Green, and C ¸a˘ gan H.S ¸ekercio˘ glu. Review: Covid-19 highlights the importance of camera traps for wildlife conservation research and management. Biological Conservation, 256:108984, 2021. ISSN 0006-3207. doi: https://doi.org/10.1016/j.biocon.2021.108984. URL https://www.sciencedirect. com/science/article/...
-
[5]
Yolov4: Optimal speed and accuracy of object detection
Alexey Bochkovskiy, Chien-Yao Wang, and Hong-Yuan Mark Liao. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934 , 2020
arXiv 2004
-
[6]
Cole Burton, Eric Neilson, Dario Moreira, Andrew Ladle, Robin Steenweg, Jason T
A. Cole Burton, Eric Neilson, Dario Moreira, Andrew Ladle, Robin Steenweg, Jason T. Fisher, Erin Bayne, and Stan Boutin. Review: Wildlife camera trapping: a review 107 and recommendations for linking surveys to ecological processes. Journal of Applied Ecology, 52(3):675–685, 2015. doi: https://doi.org/10.1111/1365-2664.12432. URL https: //besjournals.onli...
-
[7]
Banks, A Cole Burton, Caroline M
Anthony Caravaggi, Peter B. Banks, A Cole Burton, Caroline M. V. Finlay, Peter M. Haswell, Matt W. Hayward, Marcus J. Rowcliffe, and Mike D. Wood. A review of camera trapping for conservation behaviour research. Remote Sensing in Ecology and Conservation, 3(3):109–122, 2017. doi: https://doi.org/10.1002/rse2.48. URL https: //zslpublications.onlinelibrary....
doi:10.1002/rse2.48 2017
-
[8]
Christin Carl, Fiona Sch¨ onfeld, Ingolf Profft, Alisa Klamm, and Dirk Landgraf. Automated detection of european wild mammal species in camera trap images with an existing and pre-trained computer vision model. European Journal of Wildlife Research , 66, 07 2020. doi: 10.1007/s10344-020-01404-y
Show all 69 references
-
[9]
A first step towards automated species recognition from camera trap images of mammals using ai in a european temperate forest
Mateusz Choi´ nski, Mateusz Rogowski, Piotr Tynecki, Dries PJ Kuijper, Marcin Churski, and Jakub W Bubnicki. A first step towards automated species recognition from camera trap images of mammals using ai in a european temperate forest. In Computer Information Systems and Indus...
2021
-
[10]
Cole Burton
Mitchell Fennell, Christopher Beirne, and A. Cole Burton. Use of object detection in camera trap image identification: Assessing a method to rapidly and accurately classify human and animal detections for research and application in recreation ecology. Global Ecology and Conse...
2022
-
[11]
Yolox: Exceeding yolo series in 2021
Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, and Jian Sun. Yolox: Exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430 , 2021
2021 arXiv
-
[12]
Fast r-cnn
Ross Girshick. Fast r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 1440–1448, 2015
2015
-
[13]
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 580–587, 2014
2014
-
[14]
Optimising camera traps for monitoring small mammals
Alistair S Glen, Stuart Cockburn, Margaret Nichols, Jagath Ekanayake, and Bruce Warburton. Optimising camera traps for monitoring small mammals. PloS one, 8(6): e67940, 2013
2013
-
[15]
Towards automatic wild animal monitoring: Identification of animal species in camera-trap images using very deep convolutional neural networks
Alexander Gomez Villa, Augusto Salazar, and Francisco Vargas. Towards automatic wild animal monitoring: Identification of animal species in camera-trap images using very deep convolutional neural networks. Ecological Informatics, 41:24–32, 2017. ISSN 1574-9541. doi: https://do...
2017 doi
-
[16]
Improving the detection and positioning of camouflaged objects in yolov8
Tong Han, Tieyong Cao, Yunfei Zheng, Lei Chen, Yang Wang, and Bingyang Fu. Improving the detection and positioning of camouflaged objects in yolov8. Electronics, 12(20), 2023. ISSN 2079-9292. doi: 10.3390/electronics12204213. URL https://www. mdpi.com/2079-9292/12/20/4213
2023 doi
-
[17]
Long-tailed metrics and object detection in camera trap datasets
Wentong He, Ze Luo, Xinyu Tong, Xiaoyi Hu, Can Chen, and Zufei Shu. Long-tailed metrics and object detection in camera trap datasets. Applied Sciences, 13(10), 2023. ISSN 2076-3417. doi: 10.3390/app13106029. URL https://www.mdpi.com/2076-3417/ 13/10/6029
2023 doi
-
[18]
Huang, L
Z. Huang, L. Li, G. Krizek, and L. Sun. Research on traffic sign detection based on 109 improved yolov8. Journal of Computer and Communications , 11:226–232, 2023. doi: 10.4236/jcc.2023.117014
2023
-
[19]
Bgf-yolo: Enhanced yolov8 with multiscale attentional feature fusion for brain tumor detection
Ming Kang, Chee-Ming Ting, Fung Fung Ting, and Rapha¨ el C-W Phan. Bgf-yolo: Enhanced yolov8 with multiscale attentional feature fusion for brain tumor detection. arXiv preprint arXiv:2309.12585 , 2023
2023 arXiv
-
[20]
Camera traps as sensor networks for monitoring animal communities
Roland Kays, Bart Kranstauber, Patrick Jansen, Chris Carbone, Marcus Rowcliffe, Tony Fountain, and Sameer Tilak. Camera traps as sensor networks for monitoring animal communities. In 2009 IEEE 34th Conference on Local Computer Networks , pages 811–818, 2009. doi: 10.1109/LCN.2...
2009
-
[22]
Object-aware domain generalization for object detection
Wooju Lee, Dasol Hong, Hyungtae Lim, and Hyun Myung. Object-aware domain generalization for object detection. arXiv preprint arXiv:2312.12133 , 2023
2023 arXiv
-
[23]
Human vs
Scott Leorna and Todd Brinkman. Human vs. machine: Detecting wildlife in camera trap images. Ecological Informatics, 72:101876, 2022. ISSN 1574-9541. doi: https://doi. org/10.1016/j.ecoinf.2022.101876. URL https://www.sciencedirect.com/science/ article/pii/S1574954122003260
2022
-
[24]
Yolov6: A single-stage object detection framework for industrial applications
Chuyi Li, Lulu Li, Hongliang Jiang, Kaiheng Weng, Yifei Geng, Liang Li, Zaidan Ke, Qingyuan Li, Meng Cheng, Weiqiang Nie, et al. Yolov6: A single-stage object detection framework for industrial applications. arXiv preprint arXiv:2209.02976 , 2022
2022 arXiv
-
[25]
A modified yolov8 detection network for uav aerial image recognition
Yiting Li, Qingsong Fan, Haisong Huang, Zhenggong Han, and Qiang Gu. A modified yolov8 detection network for uav aerial image recognition. Drones, 7:304, 05 2023. doi: 10.3390/drones7050304. 110
2023 doi
-
[26]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedin...
2014
-
[27]
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Doll´ ar. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, pages 2980–2988, 2017
2017
-
[28]
Precise detection for dense pcb components based on modified yolov8
Qin Ling, Nor Ashidi Mat Isa, and Mohd Shahrimie Mohd Asaari. Precise detection for dense pcb components based on modified yolov8. IEEE Access, PP:1–1, 01 2023. doi: 10.1109/ACCESS.2023.3325885
2023
-
[29]
Dsw-yolov8n: A new underwater target detection algorithm based on improved yolov8n
Qiang Liu, Wei Huang, Xiaoqiu Duan, Jianghao Wei, Tao Hu, Jie Yu, and Jiahuan Huang. Dsw-yolov8n: A new underwater target detection algorithm based on improved yolov8n. Electronics, 12(18), 2023. ISSN 2079-9292. doi: 10.3390/electronics12183892. URL https://www.mdpi.com/2079-9...
2023 doi
-
[30]
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg. SSD: Single Shot MultiBox Detector , page 21–37. Springer International Publishing, 2016. ISBN 9783319464480. doi: 10.1007/ 978-3-319-46448-0 2. URL http://dx.doi.or...
2016 doi
-
[31]
Global attention mechanism: Retain information to enhance channel-spatial interactions
Yichao Liu, Zongru Shao, and Nico Hoffmann. Global attention mechanism: Retain information to enhance channel-spatial interactions. arXiv preprint arXiv:2112.05561 , 2021
2021 arXiv
-
[32]
Towards automatic detection of animals in camera-trap images
Alexander Loos, Christian Weigel, and Mona Koehler. Towards automatic detection of animals in camera-trap images. In 2018 26th European Signal Processing Conference (EUSIPCO), pages 1805–1809, 2018. doi: 10.23919/EUSIPCO.2018.8553439. 111
2018
-
[33]
H. Lou, X. Duan, J. Guo, H. Liu, J. Gu, L. Bi, and H. Chen. Dc-yolov8: Small-size object detection algorithm based on camera sensor. Electronics, 12:2323, 2023
2023
-
[34]
Improved yolov8 detection algorithm in security inspection image
Liyao Lu. Improved yolov8 detection algorithm in security inspection image. arXiv preprint arXiv:2308.06452, 2023
2023 arXiv
-
[35]
Sp-yolov8s: An improved yolov8s model for remote sensing image tiny object detection
Mingyang Ma and Huanli Pang. Sp-yolov8s: An improved yolov8s model for remote sensing image tiny object detection. Applied Sciences, 13(14), 2023. ISSN 2076-3417. doi: 10.3390/app13148161. URL https://www.mdpi.com/2076-3417/13/14/8161
2023 doi
-
[36]
Time to automate identification
Norman MacLeod, Mark Benfield, and Phil Culverhouse. Time to automate identification. Nature, 467:154–5, 09 2010. doi: 10.1038/467154a
2010 doi
-
[37]
Limitations of recreational camera traps for wildlife management and conservation research: A practitioner’s perspective
Scott Newey, Paul Davidson, Sajid Nazir, Gorry Fairhurst, Fabio Verdicchio, Robert Irvine, and Rene van der Wal. Limitations of recreational camera traps for wildlife management and conservation research: A practitioner’s perspective. Ambio, 44:624–635, 11 2015. doi: 10.1007/s...
2015 doi
-
[38]
Camera traps in animal ecology and conservation: What’s next? Camera traps in animal ecology: methods and analyses, pages 253–263, 2011
James D Nichols, Allan F O’Connell, and K Ullas Karanth. Camera traps in animal ecology and conservation: What’s next? Camera traps in animal ecology: methods and analyses, pages 253–263, 2011
2011
-
[39]
Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning
Mohammad Sadegh Norouzzadeh, Anh Nguyen, Margaret Kosmala, Alexandra Swanson, Meredith S Palmer, Craig Packer, and Jeff Clune. Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning. Proceedings of the National Academy of Scie...
2018
-
[40]
A deep active learning system for species identification and counting in camera trap images
Mohammad Sadegh Norouzzadeh, Dan Morris, Sara Beery, Neel Joshi, Nebojsa Jojic, and Jeff Clune. A deep active learning system for species identification and counting in camera trap images. Methods in Ecology and Evolution, 12(1):150–161, 2021. doi: https:// doi.org/10.1111/204...
2021 doi
-
[41]
Yolo9000: better, faster, stronger
Joseph Redmon and Ali Farhadi. Yolo9000: better, faster, stronger. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 7263–7271, 2017
2017
-
[42]
Yolov3: An incremental improvement
Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018
2018 arXiv
-
[43]
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 779–788, 2016
2016
-
[44]
Faster r-cnn: Towards real-time object detection with region proposal networks.Advances in neural information processing systems, 28, 2015
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks.Advances in neural information processing systems, 28, 2015
2015
-
[45]
Generalized intersection over union: A metric and a loss for bounding box regression
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 658–666, 2019
2019
-
[47]
Marcus Rowcliffe, Juliet Field, Samuel T
J. Marcus Rowcliffe, Juliet Field, Samuel T. Turvey, and Chris Carbone. Estimating animal density using camera traps without the need for individual recognition. Journal of Applied Ecology, 45(4):1228–1236, 2008. doi: https://doi.org/10.1111/j.1365-2664.2008. 01473.x. URL http...
2008
-
[48]
Taylor, and Stefan C
Stefan Schneider, Saul Greenberg, Graham W. Taylor, and Stefan C. Kremer. Three critical factors affecting automated image species recognition performance for camera 113 traps. Ecology and Evolution , 10(7):3503–3517, 2020. doi: https://doi.org/10.1002/ece3
2020 doi
-
[49]
Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. International Journal of Computer Vision , 128(2): 336–359, October 2019. ISSN 1573-1...
2019 doi
-
[50]
Swann, Kae Kawanishi, and Jonathan Palmer
Don E. Swann, Kae Kawanishi, and Jonathan Palmer. Evaluating Types and Features of Camera Traps in Ecological Studies: A Guide for Researchers , pages 27–43. Springer Japan, Tokyo, 2011. ISBN 978-4-431-99495-4. doi: 10.1007/978-4-431-99495-4 3. URL https://doi.org/10.1007/978-...
2011 doi
-
[51]
Tabak, Daniel Falbel, Tess Hamzeh, Ryan K
Michael A. Tabak, Daniel Falbel, Tess Hamzeh, Ryan K. Brook, John A. Goolsby, Lisa D. Zoromski, Raoul K. Boughton, Nathan P. Snow, Kurt C. VerCauteren, and Ryan S. Miller. Cameratrapdetector: Automatically detect, classify, and count animals in camera trap images using artific...
2022 doi
-
[52]
Camera trap placement for evaluating species richness, abundance, and activity
Kamakshi Tanwar, Ayan Sadhu, and Yadvendradev Jhala. Camera trap placement for evaluating species richness, abundance, and activity. Scientific Reports, 11, 11 2021. doi: 10.1038/s41598-021-02459-w
2021 doi
-
[53]
torch.nn.Linear
PyTorch Team. torch.nn.Linear. https://pytorch.org/docs/stable/generated/ torch.nn.Linear.html. Accessed: 21st Feb 2024
2024
-
[54]
A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas
Juan Terven, Diana-Margarita C´ ordova-Esparza, and Julio-Alejandro Romero-Gonz´ alez. A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas. Machine Learning and Knowledge Extraction , 5(4):1680–1716, November 114
-
[55]
M. W. Tobler, S. E. Carrillo-Percastegui, R. Leite Pitman, R. Mares, and G. Powell. An evaluation of camera traps for inventorying large- and medium-sized terrestrial rainforest mammals. Animal Conservation, 11(3):169–178, 2008. doi: https://doi.org/ 10.1111/j.1469-1795.2008.0...
2008
-
[56]
Wise-iou: Bounding box regression loss with dynamic focusing mechanism
Z Tong, Y Chen, Z Xu, and R Yu. Wise-iou: Bounding box regression loss with dynamic focusing mechanism. arXiv preprint arXiv:2301.10051 , 2023
2023 arXiv
-
[57]
Sande, T
Jasper Uijlings, K. Sande, T. Gevers, and A.W.M. Smeulders. Selective search for object recognition. International Journal of Computer Vision , 104:154–171, 09 2013. doi: 10.1007/s11263-013-0620-5
2013 doi
-
[58]
YOLOv5: You Only Look Once for Object Detection
Ultralytics. YOLOv5: You Only Look Once for Object Detection. https://github. com/ultralytics/yolov5, 2023. Accessed: October 12, 2023
2023
-
[59]
YOLOv8: You Only Look Once for Object Detection
Ultralytics. YOLOv8: You Only Look Once for Object Detection. https://github. com/ultralytics/ultralytics, 2023. Accessed: February 04, 2024
2023
-
[60]
Cspnet: A new backbone that can enhance learning capability of cnn
Chien-Yao Wang, Hong-Yuan Mark Liao, Yueh-Hua Wu, Ping-Yang Chen, Jun-Wei Hsieh, and I-Hau Yeh. Cspnet: A new backbone that can enhance learning capability of cnn. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 390–391, 2020
2020
-
[61]
Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
Chien-Yao Wang, Alexey Bochkovskiy, and Hong-Yuan Mark Liao. Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 7464–7475, 2023. 115
2023
-
[62]
Uav-yolov8: A small-object-detection model based on improved yolov8 for uav aerial photography scenarios
Gang Wang, Yanfei Chen, Pei An, Hanyu Hong, Jinghu Hu, and Tiange Huang. Uav-yolov8: A small-object-detection model based on improved yolov8 for uav aerial photography scenarios. Sensors, 23(16), 2023. ISSN 1424-8220. doi: 10.3390/s23167190. URL https://www.mdpi.com/1424-8220/...
2023 doi
-
[63]
Bl-yolov8: An improved road defect detection model based on yolov8
Xueqiu Wang, Huanbing Gao, Zemeng Jia, and Zijian Li. Bl-yolov8: An improved road defect detection model based on yolov8. Sensors, 23:8361, 10 2023. doi: 10.3390/ s23208361
2023
-
[64]
K. Xia, Z. Lv, K. Liu, et al. Global contextual attention augmented yolo with convmixer prediction heads for pcb surface defect detection. Scientific Reports, 13:9805, 2023. doi: 10.1038/s41598-023-36854-2
2023 doi
-
[65]
A lightweight yolov8 tomato detection algorithm combining feature enhancement and attention
Guoliang Yang, Jixiang Wang, Ziling Nie, Hao Yang, and Shuaiying Yu. A lightweight yolov8 tomato detection algorithm combining feature enhancement and attention. Agronomy, 13:1824, 07 2023. doi: 10.3390/agronomy13071824
2023 doi
-
[66]
Towards domain generalization in object detection
Xingxuan Zhang, Zekai Xu, Renzhe Xu, Jiashuo Liu, Peng Cui, Weitao Wan, Chong Sun, and Chen Li. Towards domain generalization in object detection. arXiv preprint arXiv:2203.14387, 2022
2022 arXiv
-
[67]
Focal and efficient iou loss for accurate bounding box regression
Yi-Fan Zhang, Weiqiang Ren, Zhang Zhang, Zhen Jia, Liang Wang, and Tieniu Tan. Focal and efficient iou loss for accurate bounding box regression. Neurocomputing, 506: 146–157, 2022
2022
-
[68]
Distance-iou loss: Faster and better learning for bounding box regression
Zhaohui Zheng, Ping Wang, Wei Liu, Jinze Li, Rongguang Ye, and Dongwei Ren. Distance-iou loss: Faster and better learning for bounding box regression. In Proceedings of the AAAI Conference on Artificial Intelligence , pages 12993–13000, 2020. 116 Appendix A Code Repository In ...
2020
-
[2022]
URL https://onlinelibrary.wiley
doi: https://doi.org/10.1111/ele.14123. URL https://onlinelibrary.wiley. com/doi/abs/10.1111/ele.14123
-
[2023]
doi: 10.3390/make5040083
ISSN 2504-4990. doi: 10.3390/make5040083. URL http://dx.doi.org/10.3390/ make5040083
-
[6147]
URL https://onlinelibrary.wiley.com/doi/abs/10.1002/ece3.6147
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.