Pith. sign in

REVIEW 2 cited by

Knowledge Distillation in YOLOX-ViT for Side-Scan Sonar Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.09313 v1 pith:GKBIPORU submitted 2024-03-14 cs.CV cs.AI

classification cs.CVcs.AI
keywords detectiondistillationknowledgeobjectyolox-vitlayermodelperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we present YOLOX-ViT, a novel object detection model, and investigate the efficacy of knowledge distillation for model size reduction without sacrificing performance. Focused on underwater robotics, our research addresses key questions about the viability of smaller models and the impact of the visual transformer layer in YOLOX. Furthermore, we introduce a new side-scan sonar image dataset, and use it to evaluate our object detector's performance. Results show that knowledge distillation effectively reduces false positives in wall detection. Additionally, the introduced visual transformer layer significantly improves object detection accuracy in the underwater environment. The source code of the knowledge distillation in the YOLOX-ViT is at https://github.com/remaro-network/KD-YOLOX-ViT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sonar-based Deep Learning in Underwater Robotics: Overview, Robustness and Challenges

    cs.RO 2024-12 conditional novelty 5.0 of 10

    A survey of sonar-based deep learning that identifies robustness, dataset scarcity, and sim-to-real gaps as the main obstacles to safe underwater autonomy.

  2. Sensing for Space Safety and Sustainability: A Deep Learning Approach with Vision Transformers

    cs.CV 2024-12 conditional novelty 4.0 of 10

    GELAN-ViT and GELAN-RepViT reach accuracy within roughly one point of YOLOv9-t while cutting reported GFLOPs by more than five, but no error bars, code, or dataset are provided.

Pith tools