REVIEW 4 major objections 6 minor 37 references
MID: A Comprehensive Shore-Based Dataset for Multi-Scale Dense Ship Occlusion and Interaction Scenarios
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read MID, a new shore-based dataset with 135,884 oriented-box ship instances, gives busy-port detection a realistic occlusion-heavy benchmark that satellite and SAR datasets lack.
desk verdict Potentially valuable shore-based OBB ship dataset, but the unexplained 15,050-to-5,673 frame selection gap must be addressed before the diversity claims are credible. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the oriented bounding box (OBB): a rotated rectangle specified by four corner coordinates, chosen over axis-aligned boxes because ships appear at arbitrary headings and dense pixel overlap makes horizontal boxes inaccurate. Equally load-bearing is the dataset's annotation and organization scheme—each image is tied to a video ID and frame ID extracted at one frame every 176 frames (roughly every 6 seconds) from 1920×1080 shore-mounted cameras, giving a temporal ordering that detection alone would not provide. Around this scheme, the paper builds a difficulty taxonomy (weather, scale, aspect ratio, occlusion degree, background, and collision type) that lets the dataset be sliced into focused test conditions, and it runs ten YOLO-family detectors with OBB heads under fixed training settings to supply reference numbers.
What would settle it
Annotate all 15,050 extracted frames, or a random sample of the 9,377 frames that MID leaves out, and compare their instance counts, occlusion rates, weather conditions, and video coverage with MID's 5,673 images; a systematic mismatch would show that MID's diversity statistics describe the chosen subset, not the captured navigation scenes.
Extended reading notes
Core claim
MID is a video-derived optical dataset whose images carry a time dimension (video ID and frame ID) and point-based oriented bounding box annotations in the form of four corner coordinates. The dataset's defining claim is that dense occlusion and interaction-rich scenes, not just clean single-ship views, are the norm in real port monitoring: 16% of instances are at least slightly occluded, 837 instances are almost fully or fully occluded, and roughly half of all instances are tiny (at most 16 pixels) while the other half are extra-large (above 256 pixels). By including these cases alongside rain, fog, lens water droplets, overexposure, and multiple camera viewpoints, the authors aim to provide a harder and more realistic training and evaluation ground than existing datasets, one that supports both supervised and semi-supervised learning and downstream tasks such as tracking, trajectory extraction and prediction, and traffic information analysis. The baseline runs of ten YOLO variants with oriented-box heads are presented as first reference results on this benchmark.
Load-bearing premise
The load-bearing premise is that the 5,673 annotated images fairly represent the 15,050 frames extracted from the 43 videos, yet the paper gives no selection or filtering procedure between the two sets.
Editorial extensions
If this is right
- Detectors trained on MID should transfer better to crowded port and narrow-channel monitoring than models trained on satellite, SAR, or single-target datasets, because the training distribution includes occluded, tiny, and overlapping ships.
- The video ID and frame ID naming makes MID usable for tracking and trajectory extraction without extra alignment, directly supporting speed estimation and ship counting.
- The graded occlusion annotations let researchers measure how detection performance degrades as occlusion increases, and provide a test set for occlusion-aware detectors.
- The extreme scale distribution—about half tiny and half extra-large instances—stresses multiscale detectors and makes the dataset a demanding benchmark for small-target detection.
- The fixed training settings and ten baseline configurations provide a reproducible comparison point for future oriented-box detectors.
Reading between the lines
- Because the paper does not say how the 5,673 annotated images were selected from the 15,050 extracted frames, the dataset's diversity statistics implicitly assume that this subset represents the full video corpus; releasing the selection rule or all frames would let users test that assumption.
- The occlusion labels include fully occluded instances, which are not visible; this opens a route to evaluating track-based re-identification or weakly supervised detection that the paper does not develop.
- All data come from one port area during 10 days in March, so generalizations to other seasons, regions, or port layouts are plausible but untested; a straightforward check is training on MID and testing on a second port's footage.
- The time dimension plus OBB annotations could support a unified detection-and-tracking benchmark with occlusion-conditioned metrics, a construction the paper leaves for future versions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces MID, a shore-based optical maritime dataset with 5,673 images and 135,884 oriented-bounding-box (OBB) annotations derived from 43 video segments recorded by port surveillance cameras. The paper documents dataset organization, analyzes diversity in terms of weather, scale, aspect ratio, background, occlusion, and collision scenarios, and reports detection baselines for ten YOLO-family variants with OBB heads. The central claim is that MID fills a gap in existing ship datasets by providing dense, occluded, multi-scale real-world maritime interactions and therefore supports detection, tracking, and trajectory-prediction research.
Significance. If the dataset is released as described, MID is a potentially valuable community resource: shore-based OBB video-derived data with temporal ordering is scarce, and the internal statistics are arithmetically consistent (the instance counts in Tables II and III sum to 135,884, and 135,884/5,673 = 23.95). The paper also provides reproducible-looking OBB conversion tooling and a public release plan. The significance is conditional, however, on resolving the undocumented gap between the 15,050 extracted frames and the 5,673 released images, and on clarifying the definitions behind the scale and occlusion statistics.
major comments (4)
- [III.A and IV.B] The paper states in Section III.A that 43 video segments, each sampled at one frame per 176 frames, yield 350 images per video and 15,050 original images, yet the released dataset and all subsequent statistics use 5,673 images. No filtering, exclusion, or subsampling procedure is described anywhere in the manuscript. This is load-bearing because the dataset's claims to reflect real-world shore-based maritime distributions rest on the representativeness of the final image set; the phrase in Section III.B.3, 'we collected as much occluded data as possible,' and the abstract's mention of 'manually supplemented annotations' suggest curation rather than random sampling. Please specify exactly how 5,673 images were obtained from 15,050 frames, report any exclusion criteria (e.g., empty frames, blur, annotation difficulty), and state whether selection was randomized and, if so, with what seed or protocol.
- [Table V] The reported recall of 0.917 for YOLOv10s-obb head is a striking outlier: every other model in the table has recall between 0.688 and 0.722, including YOLOv11s-obb, which has the same mAP50 of 84.9. No explanation or experimental note accompanies this value, and it is unlikely to be correct as reported. Because the baseline comparison is one of the paper's central evaluation claims, please verify the YOLOv10s result, report corrected numbers, and, ideally, include variance over multiple runs or seeds.
- [Table II and Section IV.B] The scale categories in Table II are labeled only as pixel thresholds (e.g., 'Tiny Instances ≤ 16 pixels'), without specifying whether the threshold refers to bounding-box area, long side, short side, or some other quantity. The resulting distribution—65,748 tiny instances and 66,984 extra-large instances, with almost no instances in the intermediate bins—is surprising for a dataset advertised as multi-scale and needs explanation. If the threshold is on area, 16 square pixels is far below a plausible annotatable ship size; if it is on side length, the units are unspecified. Please define the scale measure and discuss the apparent bimodality, which may also indicate that the 'tiny' and 'extra-large' bins are not measuring what the text implies.
- [III.B.3 and IV.F] The paper's central occlusion contribution lacks a formal definition. Section III.B.3 states that the authors 'annotate both the visible parts of the hull and the obscured sections at different visibility ratios,' but each object has a single OBB; it is unclear how a single box can encode both visible and occluded portions or how the occlusion percentage in Table III is computed. Section IV.F says 16% of the dataset contains occlusion, but Table III reports instance counts, not image counts. Please define the occlusion ratio, describe the annotation protocol for occluded targets (e.g., is the box drawn around the full extent or only the visible part?), and report occlusion statistics at both image and instance levels.
minor comments (6)
- [IV.A] The weather categories 'silty,' 'fuzzy,' and 'color-distorted' are non-standard and are not defined; please replace them with standard meteorological or visual categories or provide quantitative criteria.
- [IV.F] The sentence '16% of the dataset contains varying levels of occlusion' should say '16% of instances' if it refers to Table III, or should be recomputed at image level.
- [Table VI] The column 'Time Dimension Year' is confusing: the checkmark for Ours appears to refer to the video/frame ID naming convention rather than to temporal annotations; please clarify what is being compared.
- [V] The baseline experiments evaluate models trained and tested on MID only; a cross-dataset evaluation (e.g., fine-tune on MID and test on HRSID, HRSC2016, or SeaShip) would substantiate the claim that MID improves generalization to real-world complex scenes.
- [III.B] Annotation quality is stated to be ensured by four experienced annotators over three months, but no inter-annotator agreement, quality-control, or re-check procedure is described; a brief protocol statement would be valuable for dataset reliability.
- [References] Reference [30] for YOLO is incomplete (missing co-authors and publication details), and there are occasional formatting inconsistencies in the reference list (e.g., incomplete venue names).
Circularity Check
No circularity found: the paper constructs a dataset and reports benchmark evaluations; no result is derived from or equivalent to its inputs.
full rationale
This is a dataset-and-benchmark paper, not a derivation. The central contribution is the MID image set with OBB annotations; its construction is described in Section III (frame extraction at one frame every 176 frames, annotation by professionals) and its properties are then measured in Section IV (weather, scale, aspect ratio, background, occlusion, collision statistics). These statistics are descriptive of the released images, not predictions obtained from a fitted model, so the paper does not claim to predict any quantity from first principles. The utility claim is supported by evaluation of ten YOLO-family detectors on a train/validation/test split of the same dataset (Section V), which is standard benchmark practice rather than circularity: the baseline results are measured performance, not quantities constructed from the models' own assumptions. There are no fitted parameters masquerading as predictions, and no load-bearing self-citations; the few citations to prior datasets and detectors are external. The unexplained reduction from 15,050 extracted frames to 5,673 released images (Section III.A vs. the abstract) is a genuine reporting gap about sample selection and representativeness, but it is a completeness and validity concern, not a circularity: nothing in the paper's argument reduces by definition to its own inputs. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Occlusion degree thresholds =
10%, 20%, 50%, 90%
- Scale category thresholds =
16, 32, 96, 256 pixels
- Frame extraction interval =
1 frame per 176 frames (approx. 6 seconds)
assumptions (3)
- domain assumption Oriented bounding boxes are more accurate than horizontal bounding boxes for ship detection in complex scenes.
- domain assumption Existing ship detection datasets (HRSID, SSDD, NWPU-10) do not adequately cover dense occlusion and interaction scenarios.
- domain assumption The camera installations and selected water areas (41 square km, 43 video segments) are representative of busy port and narrow-channel navigation.
Cite this review
Pith. "Pith review of MID: A Comprehensive Shore-Based Dataset for Multi-Scale Dense Ship Occlusion and Interaction Scenarios." pith.science (2026). https://pith.science/paper/WONACGWJ
@misc{pith2026241205871,
author = {Pith},
title = {Pith review of: MID: A Comprehensive Shore-Based Dataset for Multi-Scale Dense Ship Occlusion and Interaction Scenarios},
year = {2026},
howpublished = {\url{https://pith.science/paper/WONACGWJ}},
note = {Machine review of arXiv:2412.05871}
}
read the original abstract
This paper introduces the Maritime Ship Navigation Behavior Dataset (MID), designed to address challenges in ship detection within complex maritime environments using Oriented Bounding Boxes (OBB). MID contains 5,673 images with 135,884 finely annotated target instances, supporting both supervised and semi-supervised learning. It features diverse maritime scenarios such as ship encounters under varying weather, docking maneuvers, small target clustering, and partial occlusions, filling critical gaps in datasets like HRSID, SSDD, and NWPU-10. MID's images are sourced from high-definition video clips of real-world navigation across 43 water areas, with varied weather and lighting conditions (e.g., rain, fog). Manually curated annotations enhance the dataset's variety, ensuring its applicability to real-world demands in busy ports and dense maritime regions. This diversity equips models trained on MID to better handle complex, dynamic environments, supporting advancements in maritime situational awareness. To validate MID's utility, we evaluated 10 detection algorithms, providing an in-depth analysis of the dataset, detection results from various models, and a comparative study of baseline algorithms, with a focus on handling occlusions and dense target clusters. The results highlight MID's potential to drive innovation in intelligent maritime traffic monitoring and autonomous navigation systems. The dataset will be made publicly available at https://github.com/VirtualNew/MID_DataSet.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
An outlook on the future marine traffic management system for autonomous ships,
M. Martelli, A. Virdis, A. Gotta, P. Cassar `a, and M. Di Summa, “An outlook on the future marine traffic management system for autonomous ships,” IEEE Access , vol. 9, pp. 157 316–157 328, 2021
work page 2021
-
[2]
H. J. Peters, “Developments in global seatrade and container shipping markets: their effects on the port industry and private sector involve- ment,” Int. J. Marit. Econ. , vol. 3, no. 1, pp. 3–26, 2001
work page 2001
-
[3]
Internet of things for smart ports: Technologies and challenges,
Y . Yang, M. Zhong, H. Yao, F. Yu, X. Fu, and O. Postolache, “Internet of things for smart ports: Technologies and challenges,” IEEE Instrum. Meas. Mag. , vol. 21, no. 1, pp. 34–43, 2018
work page 2018
-
[4]
X. Xin, Z. Yang, K. Liu, J. Zhang, and X. Wu, “Multi-stage and multi- topology analysis of ship traffic complexity for probabilistic collision detection,” Expert Syst. Appl. , vol. 213, p. 118890, 2023
work page 2023
-
[5]
A sidelobe-aware small ship detection network for synthetic aperture radar imagery,
Y . Zhou, H. Liu, F. Ma, Z. Pan, and F. Zhang, “A sidelobe-aware small ship detection network for synthetic aperture radar imagery,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–16, 2023
work page 2023
-
[6]
Sensors and ai techniques for situational awareness in au- tonomous ships: A review,
S. Thombre, Z. Zhao, H. Ramm-Schmidt, J. M. V . Garc´ıa, T. Malkam¨aki, S. Nikolskiy, T. Hammarberg, H. Nuortie, M. Z. H. Bhuiyan, S. S ¨arkk¨a et al. , “Sensors and ai techniques for situational awareness in au- tonomous ships: A review,” IEEE trans. Intell. Transp. Syst. , vol. 23, no. 1, pp. 64–83, 2020
work page 2020
-
[7]
The ocean-going autonomous ship—challenges and threats,
A. Felski and K. Zwolak, “The ocean-going autonomous ship—challenges and threats,” J. Mar . Sci. Eng. , vol. 8, no. 1, p. 41, 2020
work page 2020
-
[8]
Weather-aware object detection method for maritime surveillance systems,
M. Chen, J. Sun, K. Aida, and A. Takefusa, “Weather-aware object detection method for maritime surveillance systems,” Future Gener . Comp. Sy., vol. 151, pp. 111–123, 2024
work page 2024
Show all 37 references
-
[9]
Object detection in a maritime environment: Performance evaluation of background subtraction methods,
D. K. Prasad, C. K. Prasath, D. Rajan, L. Rachmawati, E. Rajabally, and C. Quek, “Object detection in a maritime environment: Performance evaluation of background subtraction methods,” IEEE trans. Intell. Transp. Syst., vol. 20, no. 5, pp. 1787–1802, 2018
2018
-
[10]
Ship detection in high-resolution optical imagery based on anomaly detector and local shape feature,
Z. Shi, X. Yu, Z. Jiang, and B. Li, “Ship detection in high-resolution optical imagery based on anomaly detector and local shape feature,” IEEE Trans. Geosci. Remote Sens. , vol. 52, no. 8, pp. 4511–4523, 2013
2013
-
[11]
How big data enriches maritime research–a critical review of automatic identification system (ais) data applications,
D. Yang, L. Wu, S. Wang, H. Jia, and K. X. Li, “How big data enriches maritime research–a critical review of automatic identification system (ais) data applications,” Transp. Rev., vol. 39, no. 6, pp. 755–773, 2019
2019
-
[12]
Ship detection with high-resolution hf skywave radar,
J. Barnum, “Ship detection with high-resolution hf skywave radar,” IEEE J. OCEANIC ENG. , vol. 11, no. 2, pp. 196–209, 1986
1986
-
[13]
Conservation science and policy applications of the marine vessel automatic identification system (ais)—a review,
M. Robards, G. Silber, J. Adams, J. Arroyo, D. Lorenzini, K. Schwehr, and J. Amos, “Conservation science and policy applications of the marine vessel automatic identification system (ais)—a review,” B. MAR. SCI., vol. 92, no. 1, pp. 75–103, 2016
2016
-
[14]
Ship de- tection with spectral analysis of synthetic aperture radar: A comparison of new and well-known algorithms,
A. Marino, M. J. Sanjuan-Ferrer, I. Hajnsek, and K. Ouchi, “Ship de- tection with spectral analysis of synthetic aperture radar: A comparison of new and well-known algorithms,” Remote Sens. , vol. 7, no. 5, pp. 5416–5439, 2015
2015
-
[15]
Ship detection for visual maritime surveillance from non-stationary platforms,
Y . Zhang, Q.-Z. Li, and F.-N. Zang, “Ship detection for visual maritime surveillance from non-stationary platforms,” OCEAN ENG. , vol. 141, pp. 53–63, 2017
2017
-
[16]
Research of target detection and classification techniques using millimeter-wave radar and vision sensors,
Z. Wang, X. Miao, Z. Huang, and H. Luo, “Research of target detection and classification techniques using millimeter-wave radar and vision sensors,” Remote Sens. , vol. 13, no. 6, p. 1064, 2021
2021
-
[17]
Dataset and benchmark for ship detection in complex optical remote sensing image,
J. Hu, X. Zhi, T. Shi, J. Wang, Y . Li, and X. Sun, “Dataset and benchmark for ship detection in complex optical remote sensing image,” IEEE Trans. Geosci. Remote Sens. , 2024
2024
-
[18]
Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,
S. Wei, X. Zeng, Q. Qu, M. Wang, H. Su, and J. Shi, “Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,” IEEE Access , vol. 8, pp. 120 234–120 254, 2020
2020
-
[19]
Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,
T. Zhang, X. Zhang, J. Li, X. Xu, B. Wang, X. Zhan, Y . Xu, X. Ke, T. Zeng, H. Su et al., “Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,” Remote Sens., vol. 13, no. 18, p. 3690, 2021
2021
-
[20]
Object detection and instance segmentation in remote sensing imagery based on precise mask r-cnn,
H. Su, S. Wei, M. Yan, C. Wang, J. Shi, and X. Zhang, “Object detection and instance segmentation in remote sensing imagery based on precise mask r-cnn,” in Proc. IEEE Int. Geosci. Remote Sens. Symp. IEEE, 2019, pp. 1454–1457
2019
-
[21]
Hq- isnet: High-quality instance segmentation for remote sensing imagery,
H. Su, S. Wei, S. Liu, J. Liang, C. Wang, J. Shi, and X. Zhang, “Hq- isnet: High-quality instance segmentation for remote sensing imagery,” Remote Sens. , vol. 12, no. 6, p. 989, 2020
2020
-
[22]
A high resolution optical satellite image dataset for ship recognition and some new baselines,
Z. Liu, L. Yuan, L. Weng, and Y . Yang, “A high resolution optical satellite image dataset for ship recognition and some new baselines,” in Proc. Int. Conf. Pattern Recognit. Appl. Methods , vol. 2. SciTePress, 2017, pp. 324–331
2017
-
[23]
Shiprsimagenet: A large-scale fine-grained dataset for ship detection in high-resolution optical remote sensing images,
Z. Zhang, L. Zhang, Y . Wang, P. Feng, and R. He, “Shiprsimagenet: A large-scale fine-grained dataset for ship detection in high-resolution optical remote sensing images,” EEE J. Sel. Top Appl. Earth Obs. Remote Sens., vol. 14, pp. 8458–8472, 2021
2021
-
[24]
Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images,
Y . Han, X. Yang, T. Pu, and Z. Peng, “Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1–18, 2021
2021
-
[25]
Seaships: A large-scale precisely annotated dataset for ship detection,
Z. Shao, W. Wu, Z. Wang, W. Du, and C. Li, “Seaships: A large-scale precisely annotated dataset for ship detection,” IEEE Trans. Multimedia., vol. 20, no. 10, pp. 2593–2604, 2018
2018
-
[26]
A discriminatively trained, multiscale, deformable part model,
P. Felzenszwalb, D. McAllester, and D. Ramanan, “A discriminatively trained, multiscale, deformable part model,” in Proc. IEEE Conf. Com- put. Vision Pattern Recognit. Ieee, 2008, pp. 1–8
2008
-
[27]
Rich feature hierarchies for accurate object detection and semantic segmentation,
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proc. IEEE Conf. Comput. Vision Pattern Recognit. , 2014, pp. 580–587
2014
- [28]
-
[29]
Faster r-cnn: Towards real-time object detection with region proposal networks,
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 6, pp. 1137–1149, 2016
2016
-
[30]
You only look once: Unified, real-time object detection,
J. Redmon, “You only look once: Unified, real-time object detection,” in Proc. IEEE Conf. Comput. Vision Pattern Recognit. , 2016
2016
-
[31]
Ssd: Single shot multibox detector,
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y . Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in Proc. European Conf. Comput. Vision . Springer, 2016, pp. 21–37
2016
-
[32]
Focal loss for dense object detection,
T.-Y . Ross and G. Doll ´ar, “Focal loss for dense object detection,” in Proc. IEEE Conf. Comput. Vision Pattern Recognit. , 2017, pp. 2980– 2988
2017
-
[33]
End-to-end object detection with transformers,
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proc. European Conf. Comput. Vision . Springer, 2020, pp. 213–229
2020
-
[34]
Efficientdet: Scalable and efficient object detection,
M. Tan, R. Pang, and Q. V . Le, “Efficientdet: Scalable and efficient object detection,” in Proc. IEEE Conf. Comput. Vision Pattern Recognit. , 2020, pp. 10 781–10 790
2020
-
[35]
Yolov6: A single-stage object detection framework for industrial applications,
C. Li, L. Li, H. Jiang, K. Weng, Y . Geng, L. Li, Z. Ke, Q. Li, M. Cheng, W. Nie et al. , “Yolov6: A single-stage object detection framework for industrial applications,” arXiv preprint arXiv:2209.02976 , 2022
2022 arXiv
-
[36]
Yolov9: Learning what you want to learn using programmable gradient information,
C.-Y . Wang, I.-H. Yeh, and H.-Y . Mark Liao, “Yolov9: Learning what you want to learn using programmable gradient information,” in Proc. European Conf. Comput. Vision . Springer, 2025, pp. 1–21
2025
-
[37]
Yolov10: Real-time end-to-end object detection,
A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, J. Han, and G. Ding, “Yolov10: Real-time end-to-end object detection,” arXiv preprint arXiv:2405.14458, 2024
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.