Pith. sign in

REVIEW 2 major objections 6 minor 180 references

Imbalance Problems in Object Detection: A Review

T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This review claims that the performance of deep object detectors is governed by eight distinct imbalance problems, organized into four categories, and offers a taxonomy that maps each problem to its solutions and open issues.

desk verdict A genuinely useful survey whose four-category, eight-problem taxonomy will get cited, even though the paper's own open-issues sections quietly introduce more imbalance types than the taxonomy promises. read the letter →

arxiv 1909.00169 v3 pith:ZFCIICIP submitted 2019-08-31 cs.CV

classification cs.CV
keywords objectdetectionimbalanceproblemstaxonomyclassscalespatialobjectivedeeplearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the many ways object detectors underperform can be organized into eight specific imbalance problems, grouped into four categories: class imbalance, scale imbalance, spatial imbalance, and objective imbalance. An imbalance exists when the distribution of some input property—class counts, object scales, bounding-box locations, IoUs, or loss contributions—hurts detection performance. If this taxonomy is right, researchers gain a shared map: a problem to name, a set of known remedies to compare, and a list of gaps to attack. The review also argues that these problems interact, so treating any one in isolation may be insufficient.

What carries the argument

The taxonomy itself, organized around the 'related input property' of each imbalance, is the central object. The definitional move is that an imbalance problem exists exactly when the distribution of that property affects performance. This converts scattered observations into a map that specifies where in the pipeline each problem arises and which family of solutions applies.

What would settle it

Show that a distributional bias not among the eight—for example object orientation, annotation noise, or temporal bias in video—changes detection accuracy when all eight listed distributions are controlled. A measured performance drop tied to that bias would break the taxonomy's completeness claim; the paper itself flags orientation imbalance as unexplored.

Watch

Extended reading notes

Core claim

The paper's central discovery is a problem-based taxonomy of imbalance in deep object detection. It identifies eight problems: foreground-background class imbalance, foreground-foreground class imbalance, object/box-level scale imbalance, feature-level imbalance, imbalance in regression loss, IoU distribution imbalance, object location imbalance, and objective imbalance. Each is tied to a specific input property and located at a stage of the training pipeline. The paper argues that existing methods—hard and soft sampling, pyramid architectures, regression-loss redesigns, cascades, and task weighting—are best understood as responses to individual entries on this list, and that unsolved issues become visible once the list is explicit.

Load-bearing premise

The taxonomy's completeness: the paper assumes the eight listed problems cover every distributional bias that affects object-detection performance, but it gives no independent rule for deciding what counts as an imbalance beyond 'it affects performance'.

Editorial extensions

If this is right

  • A researcher facing a detection failure can use the taxonomy to name the imbalance and immediately see the solution families already tried for it.
  • The review's open issues—quantifying imbalance, agreeing on positive/negative labeling, building a unified approach, and studying bottom-up detectors—become concrete research targets.
  • Methods from image classification and metric learning, such as self-paced hardness and class-balanced weighting, are identified as transferable to object-detection imbalances.
  • Because the problems interact, a small geometric change in a bounding box can shift a sample across class, scale, spatial, and objective imbalance categories; fixes must account for this.
  • Bottom-up detectors are flagged as an under-explored area where known imbalance remedies and new imbalance-specific issues both need study.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the taxonomy is complete, it gives a diagnostic checklist: a detector's remaining error can be attributed to one of the eight distributions, and a benchmark audit measuring these eight distributions could predict which fix will help.
  • The paper's own examples suggest 'balanced' is not always optimal—OHEM likes a right-skewed IoU distribution and prime samples favor high IoUs—so defining the desired distribution for each property may be more fruitful than aiming for uniformity.
  • A testable prediction follows from the interaction argument: a composite method that explicitly balances all eight distributions at once should beat stacking individual fixes, since separate fixes may trade one imbalance for another.
  • The taxonomy invites extension: temporal bias in video, object orientation, and annotation noise are candidates the paper flags but leaves outside the eight; showing any of these independently hurts performance would expand or revise the map.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This manuscript presents a review of imbalance problems in deep-learning object detection. It introduces a taxonomy of eight imbalance problems grouped into four categories (class, scale, spatial, and objective), reviews solution methods for each problem, provides comparative summaries, identifies open issues, discusses imbalance in related domains, and maintains a living online bibliography. The central claim is that the taxonomy provides a comprehensive organization of the distributional biases that affect detection performance, and that the review can serve as a map for future research.

Significance. The review is valuable and timely: it consolidates a scattered literature and gives researchers a shared vocabulary and structure. Its strengths include original descriptive statistics on datasets and trained detectors (Figures 4, 5, 7, 12, 13, and 14), critical comparative tables (e.g., Tables 3, 4, and 5), a broad coverage of solution families, and a useful living webpage for tracking new work. The taxonomy, while not rigorously proven to be exhaustive, is a practical organizing device that goes beyond prior surveys that focus almost exclusively on class imbalance. If the completeness caveats are addressed, this will be a useful reference for the object-detection community.

major comments (2)
  1. [Section 3 and Table 1 vs. Sections 6.5.4-6.5.6 and 9.3] The paper defines an imbalance problem broadly as a performance-affecting distributional bias (Section 1, where the definition is given) and describes Table 1 as a complete taxonomy (Section 3, paragraph 2). However, Sections 6.5.4, 6.5.5, and 6.5.6 introduce 'Relative Spatial Distribution Imbalance,' 'Imbalance in Overlapping BBs,' and 'Orientation Imbalance' as open issues, and Section 9.3 discusses labeling ambiguity and noise. These are imbalance problems by the paper's own definition but are not entries in Table 1. The manuscript should either incorporate these into the taxonomy or explicitly define a scope condition (for example, 'imbalance problems that have a substantial dedicated solution literature') that excludes them, and adjust the 'complete taxonomy' and 'only one of eight' wording in Sections 1 and 3 accordingly.
  2. [Section 3 and Table 1 vs. Section 9.1 and Figure 18] The taxonomy is presented as grouping problems into four main categories, but the paper's own example in Figure 18 shows that a single bounding-box shift changes class imbalance, scale imbalance, spatial imbalance, and objective imbalance simultaneously. This is not a fatal flaw, but it should be acknowledged in Section 3: the taxonomy is a set of analytical viewpoints rather than a partition into disjoint problem classes. Without this clarification, the phrases 'complete taxonomy' and 'eight different imbalance problems' overstate the disjointness of the categories and invite the kind of counterexample that the paper itself provides.
minor comments (6)
  1. [Section 4.3] The relative improvement of AP Loss over Focal Loss is stated as 3.9% with values 33.9 to 35.0 mAP; these numbers imply a relative improvement of about 3.2%. Please correct the arithmetic or clarify the baseline.
  2. [Figures 4, 5, 7, 12, 13, and 14] The new descriptive statistics are a valuable contribution, but the descriptions do not fully specify the preprocessing and computational details (for example, exact dataset splits, anchor configurations, normalization conventions, and model versions). Please provide a short methodology note or a reproducibility footnote.
  3. [Section 6.2] The analysis in Figure 12, including the statement that 'it is better off without applying regression' for high-IoU bins, is based on a single converged RetinaNet model and should be presented as an illustrative case study rather than a general empirical result.
  4. [Section 4.2.2] There is a typo: 'Oksuz et a. [66]' should be 'Oksuz et al. [66]'.
  5. [Section 8.2] There is a typo: 'during traning' should be 'during training'.
  6. [Section 5.2] The definition of feature-level imbalance would be clearer if the paper formalized what is meant by 'contribution of the feature layer,' since the current wording conflates feature semantics, spatial resolution, and gradient magnitude in a single phrase.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's taxonomy is an organizational map of the literature, and its self-citations are not load-bearing.

full rationale

This is a literature review whose central product is a taxonomy of eight imbalance problems grouped into four categories (Table 1). The taxonomy is not derived from a first-principles model; it is an organizational scheme grounded in the reviewed literature. The working definition of an imbalance problem ('when the distribution regarding that property affects the performance') is used to select and organize material, not as an equation from which the eight problems are deduced. No prediction or fitted parameter is presented, so no step reduces to its own input by construction. The paper does cite the authors' own prior works — pRoI Generator [66] and LRP [102] — for example as evidence for the 'Relative Spatial Distribution Imbalance' open issue and for OFB sampling in Section 4.2.2. These citations are not load-bearing for the taxonomy's construction: they are used as published, externally checkable empirical results, and the paper explicitly marks the performance impact of the relative spatial imbalance as still uninvestigated. The later open issues (relative spatial, overlapping-BB, orientation) are explicitly presented as open problems rather than as additional taxonomy entries, so the completeness of the eight-problem list is a scope question, not a circularity. Consequently, no specific circular step can be quoted, and the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper's central contribution is a taxonomy of known problems, so it introduces no free parameters or invented physical entities. The load-bearing assumptions are the definition of imbalance and the pipeline-based decomposition that bounds the taxonomy's scope. Both are explicit and reasonable for a review, but they are assumptions rather than derived results.

assumptions (2)
  • domain assumption An imbalance problem with respect to an input property occurs when the distribution regarding that property affects the performance.
    This definition is the foundation of the taxonomy (Section 1.1 and Section 3). It is broad and partially circular: any performance-affecting distributional bias qualifies, so the set of imbalance problems is not independently bounded.
  • domain assumption The object detection training pipeline can be decomposed into feature extraction, detection, and bounding-box matching/labeling/sampling phases, and all imbalances arise within these phases.
    The taxonomy is organized around this pipeline (Figure 1 and Section 3). The assumption is that no relevant imbalance arises outside these phases, which is a modeling choice that supports the taxonomy's structure.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Imbalance Problems in Object Detection: A Review." pith.science (2026). https://pith.science/paper/ZFCIICIP

@misc{pith2026190900169,
  author       = {Pith},
  title        = {Pith review of: Imbalance Problems in Object Detection: A Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZFCIICIP}},
  note         = {Machine review of arXiv:1909.00169}
}
read the original abstract

In this paper, we present a comprehensive review of the imbalance problems in object detection. To analyze the problems in a systematic manner, we introduce a problem-based taxonomy. Following this taxonomy, we discuss each problem in depth and present a unifying yet critical perspective on the solutions in the literature. In addition, we identify major open issues regarding the existing imbalance problems as well as imbalance problems that have not been discussed before. Moreover, in order to keep our review up to date, we provide an accompanying webpage which catalogs papers addressing imbalance problems, according to our problem-based taxonomy. Researchers can track newer studies on this webpage available at: https://github.com/kemaloksuz/ObjectDetectionImbalance .

Figures

Figures reproduced from arXiv: 1909.00169 by the authors.

Figure 1
Figure 1. (a) The common training pipeline of a generic detection network. The pipeline has 3 phases (i.e. feature extraction, detection and BB matching, labeling and sampling) represented by different background colors. (b) Illustration of an example imbalance problem from each category for object detection through the training pipeline. Background colors specify at which phase an imbalance problem occurs. Feature Extraction… view at source ↗
Figure 2
Figure 2. Problem based categorization of the methods used for imbalance problems. Note that a work may appear at multiple [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Number of papers per imbalance problem category [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (15 more)
Figure 4
Figure 4. Figure 4: Illustration of the class imbalance problems. The numbers of RetinaNet [22] anchors on MS-COCO [90] are plotted [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Some statistics of common datasets (training sets). For readability, the [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Illustration of batch-level class imbalance. [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: Imbalance in scales of the objects in common datasets: the distributions of BB width [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: An illustration and comparison of the solutions [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: Feature-Level imbalance is illustrated on the FPN [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: High-level diagrams of the methods designed for feature-level imbalance. [PITH_FULL_IMAGE:figures/full_fig_p016_10.png]
Figure 11
Figure 11. Figure 11: An illustration of imbalance in regression loss. Blue [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]
Figure 12
Figure 12. Figure 12: The IoU distribution of the positive anchors for a converged RetinaNet [22] with ResNet-50 [93] on MS COCO [90] [PITH_FULL_IMAGE:figures/full_fig_p021_12.png]
Figure 13
Figure 13. Figure 13: Distribution of the centers of the objects in the common datasets over the normalized image. [PITH_FULL_IMAGE:figures/full_fig_p022_13.png]
Figure 15
Figure 15. Figure 15: An illustration of the imbalance in overlapping BBs. [PITH_FULL_IMAGE:figures/full_fig_p023_15.png]
Figure 16
Figure 16. Figure 16: (a) Randomly sampled 32 positive RoIs using the pRoI Generator [66]. (b) Average classification and regres￾sion losses of these RoIs at the initialization of the object detector for MS COCO dataset [90] with 80 classes. We use cross entropy for the classification task…
Figure 17
Figure 17. Figure 17: We observe that the classification loss decreases [PITH_FULL_IMAGE:figures/full_fig_p025_17.png]
Figure 18
Figure 18. Figure 18: An example suggesting the necessity of considering [PITH_FULL_IMAGE:figures/full_fig_p027_18.png]
Figure 19
Figure 19. Figure 19: An illustration on ambiguities resulting from la [PITH_FULL_IMAGE:figures/full_fig_p029_19.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

180 extracted references · 67 canonical work pages

  1. [1]

    A closer look at faster r-cnn for vehicle detection,

    Q. Fan, L. Brown, and J. Smith, “A closer look at faster r-cnn for vehicle detection,” in IEEE Intelligent Vehicles Symposium, 2016

  2. [2]

    Fore- ground gating and background refining network for surveillance object detection,

    Z. Fu, Y. Chen, H. Yong, R. Jiang, L. Zhang, and X. Hua, “Fore- ground gating and background refining network for surveillance object detection,” IEEE Transactions on Image Processing-Accepted , 2019

  3. [3]

    Are we ready for au- tonomous driving? the kitti vision benchmark suite,

    A. Geiger, P . Lenz, and R. Urtasun, “Are we ready for au- tonomous driving? the kitti vision benchmark suite,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2012

  4. [4]

    Hybridnet: A fast vehicle detection system for au- tonomous driving,

    X. Dai, “Hybridnet: A fast vehicle detection system for au- tonomous driving,” Signal Processing: Image Communication , vol. 70, pp. 79 – 88, 2019

  5. [5]

    Retina u-net: Embarrassingly simple exploitation of segmentation supervision for medical object detection,

    P . F. Jaeger, S. A. A. Kohl, S. Bickelhaupt, F. Isensee, T. A. Kuder, H. Schlemmer, and K. H. Maier-Hein, “Retina u-net: Embarrassingly simple exploitation of segmentation supervision for medical object detection,” arXiv, vol. 1811.08661, 2018

  6. [6]

    Liver lesion detection from weakly-labeled multi-phase ct volumes with a grouped single shot multibox detector,

    S.-g. Lee, J. S. Bae, H. Kim, J. H. Kim, and S. Yoon, “Liver lesion detection from weakly-labeled multi-phase ct volumes with a grouped single shot multibox detector,” in Medical Image Computing and Computer Assisted Intervention (MICCAI) , A. F. Frangi, J. A. Schnabel, C. Davatzikos, C. Alberola-López, and G. Fichtinger, Eds., 2018

  7. [7]

    Bb8: A scalable, accurate, robust to partial occlusion method for predicting the 3d poses of challenging objects without using depth,

    M. Rad and V . Lepetit, “Bb8: A scalable, accurate, robust to partial occlusion method for predicting the 3d poses of challenging objects without using depth,” in The IEEE International Conference on Computer Vision (ICCV), 2017

  8. [8]

    Ssd- 6d: Making rgb-based 3d detection and 6d pose estimation great again,

    W. Kehl, F. Manhardt, F. Tombari, S. Ilic, and N. Navab, “Ssd- 6d: Making rgb-based 3d detection and 6d pose estimation great again,” in The IEEE International Conference on Computer Vision (ICCV), 2017. OKSUZ et al.: IMBALANCE PROBLEMS IN OBJECT DETECTION: A REVIEW 30

Show all 180 references
  1. [9]

    Real-Time Seamless Single Shot 6D Object Pose Prediction,

    B. Tekin, S. N. Sinha, and P . Fua, “Real-Time Seamless Single Shot 6D Object Pose Prediction,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  2. [10]

    Bop: Benchmark for 6d object pose estimation,

    T. Hodan, F. Michel, E. Brachmann, W. Kehl, A. GlentBuch, D. Kraft, B. Drost, J. Vidal, S. Ihrke, X. Zabulis, C. Sahin, F. Manhardt, F. Tombari, T.-K. Kim, J. Matas, and C. Rother, “Bop: Benchmark for 6d object pose estimation,” in The European Conference on Computer Vision (E...

  3. [11]

    Vision-based robotic grasping from object localization, pose estimation, grasp detection to motion planning: A review,

    G. Du, K. Wang, and S. Lian, “Vision-based robotic grasping from object localization, pose estimation, grasp detection to motion planning: A review,” arXiv, vol. 1905.06658, 2019

  4. [12]

    Cosmo: Contextualized scene modeling with boltzmann machines,

    I. Bozcan and S. Kalkan, “Cosmo: Contextualized scene modeling with boltzmann machines,” Robotics and Autonomous Systems, vol. 113, pp. 132–148, 2019

  5. [13]

    Object detection with discriminatively trained part- based models,

    P . F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ra- manan, “Object detection with discriminatively trained part- based models,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 32, no. 9, pp. 1627–1645, 2010

  6. [14]

    Imagenet classifi- cation with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classifi- cation with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (NIPS), 2012

  7. [15]

    YOLO9000: Better, faster, stronger,

    J. Redmon and A. Farhadi, “YOLO9000: Better, faster, stronger,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017

  8. [16]

    Rich feature hierarchies for accurate object detection and semantic segmen- tation,

    R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmen- tation,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014

  9. [17]

    Fast R-CNN,

    R. Girshick, “Fast R-CNN,” in The IEEE International Conference on Computer Vision (ICCV), 2015

  10. [18]

    R-FCN: Object detection via region-based fully convolutional networks,

    J. Dai, Y. Li, K. He, and J. Sun, “R-FCN: Object detection via region-based fully convolutional networks,” inAdvances in Neural Information Processing Systems (NIPS), 2016

  11. [19]

    SSD: single shot multibox detector,

    W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. E. Reed, C. Fu, and A. C. Berg, “SSD: single shot multibox detector,” in The European Conference on Computer Vision (ECCV), 2016

  12. [20]

    You only look once: Unified, real-time object detection,

    J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016

  13. [21]

    Faster R-CNN: Towards real-time object detection with region proposal networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 39, no. 6, pp. 1137–1149, 2017

  14. [22]

    Focal loss for dense object detection,

    T. Lin, P . Goyal, R. B. Girshick, K. He, and P . Dollár, “Focal loss for dense object detection,” in The IEEE International Conference on Computer Vision (ICCV), 2017

  15. [23]

    Cornernet: Detecting objects as paired keypoints,

    H. Law and J. Deng, “Cornernet: Detecting objects as paired keypoints,” in The European Conference on Computer Vision (ECCV), 2018

  16. [24]

    Training region- based object detectors with online hard example mining,

    A. Shrivastava, A. Gupta, and R. Girshick, “Training region- based object detectors with online hard example mining,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016

  17. [25]

    Factors in finetun- ing deep model for object detection with long-tail distribution,

    W. Ouyang, X. Wang, C. Zhang, and X. Yang, “Factors in finetun- ing deep model for object detection with long-tail distribution,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016

  18. [26]

    Feature pyramid networks for object detection,

    T. Lin, P . Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie, “Feature pyramid networks for object detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017

  19. [27]

    An analysis of scale invariance in object detection - snip,

    B. Singh and L. S. Davis, “An analysis of scale invariance in object detection - snip,” in The Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  20. [28]

    Sniper: Efficient multi-scale training,

    B. Singh, M. Najibi, and L. S. Davis, “Sniper: Efficient multi-scale training,” in Advances in Neural Information Processing Systems (NIPS), 2018

  21. [29]

    Libra R-CNN: Towards balanced learning for object detection,

    J. Pang, K. Chen, J. Shi, H. Feng, W. Ouyang, and D. Lin, “Libra R-CNN: Towards balanced learning for object detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  22. [30]

    Prime Sample Attention in Object Detection,

    Y. Cao, K. Chen, C. C. Loy, and D. Lin, “Prime Sample Attention in Object Detection,” arXiv, vol. 1904.04821, 2019

  23. [31]

    Deep learning for generic object detection: A survey,

    L. Liu, W. Ouyang, X. Wang, P . W. Fieguth, J. Chen, X. Liu, and M. Pietikäinen, “Deep learning for generic object detection: A survey,” arXiv, vol. 1809.02165, 2018

  24. [32]

    Object detection in 20 years: A survey,

    Z. Zou, Z. Shi, Y. Guo, and J. Ye, “Object detection in 20 years: A survey,” arXiv, vol. 1905.05055, 2018

  25. [33]

    Recent advances in ob- ject detection in the age of deep convolutional neural networks,

    S. Agarwal, J. O. D. Terrail, and F. Jurie, “Recent advances in ob- ject detection in the age of deep convolutional neural networks,” arXiv, vol. 1809.03193, 2018

  26. [34]

    On-road vehicle detection: a review,

    Zehang Sun, G. Bebis, and R. Miller, “On-road vehicle detection: a review,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TP AMI), vol. 28, no. 5, pp. 694–711, 2006

  27. [35]

    Pedestrian detec- tion: An evaluation of the state of the art,

    P . Dollar, C. Wojek, B. Schiele, and P . Perona, “Pedestrian detec- tion: An evaluation of the state of the art,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TP AMI) , vol. 34, no. 4, pp. 743–761, 2012

  28. [36]

    A survey on face detection in the wild,

    S. Zafeiriou, C. Zhang, and Z. Zhang, “A survey on face detection in the wild,” Computer Vision and Image Understanding , vol. 138, pp. 1–24, 2015

  29. [37]

    Text detection and recognition in imagery: A survey,

    Q. Ye and D. Doermann, “Text detection and recognition in imagery: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 37, no. 7, pp. 1480–1500, 2015

  30. [38]

    Text detection, tracking and recognition in video: A comprehensive survey,

    X. Yin, Z. Zuo, S. Tian, and C. Liu, “Text detection, tracking and recognition in video: A comprehensive survey,”IEEE Transactions on Image Processing, vol. 25, no. 6, pp. 2752–2773, 2016

  31. [39]

    A survey on deep learning in medical image analysis,

    G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. A. van der Laak, B. van Ginneken, and C. I. Sánchez, “A survey on deep learning in medical image analysis,” Medical Image Analysis, vol. 42, pp. 60 – 88, 2017

  32. [40]

    Learning from imbalanced data: open challenges and future directions,

    B. Krawczyk, “Learning from imbalanced data: open challenges and future directions,” Progress in Artificial Intelligence , vol. 5, no. 4, pp. 221–232, 2016

  33. [41]

    A survey on addressing high-class imbalance in big data,

    J. L. Leevy, T. M. Khoshgoftaar, R. A. Bauder, and N. Seliya, “A survey on addressing high-class imbalance in big data,” Journal of Big Data, vol. 5, no. 42, 2018

  34. [42]

    Fernández, S

    A. Fernández, S. García, M. Galar, R. Prati, B. Krawczyk, and F. Herrera, Learning from Imbalanced Data Sets. Springer Interna- tional Publishing, 2018

  35. [43]

    Survey on deep learning with class imbalance,

    J. M. Johnson, Khoshgoftaar, and T. M., “Survey on deep learning with class imbalance,” Journal of Big Data, vol. 6, no. 21, 2019

  36. [44]

    Selective search for object recognition,

    J. R. R. Uijlings, K. E. A. van de Sande, T. Gevers, and A. W. M. Smeulders, “Selective search for object recognition,” International Journal of Computer Vision, vol. 104, no. 2, pp. 154–171, 2013

  37. [45]

    Edge boxes: Locating object propos- als from edges,

    C. L. Zitnick and P . Dollár, “Edge boxes: Locating object propos- als from edges,” in The European Conference on Computer Vision (ECCV), 2014

  38. [46]

    DSSD: Deconvolutional single shot detector,

    C.-Y. Fu, W. Liu, A. Ranga, A. Tyagi, and A. C. Berg, “DSSD: Deconvolutional single shot detector,” arXiv, vol. 1701.06659, 2017

  39. [47]

    Yolov3: An incremental improve- ment,

    J. Redmon and A. Farhadi, “Yolov3: An incremental improve- ment,” arXiv, vol. 1804.02767, 2018

  40. [48]

    Centernet: Keypoint triplets for object detection,

    K. Duan, S. Bai, L. Xie, H. Qi, Q. Huang, and Q. Tian, “Centernet: Keypoint triplets for object detection,” in The IEEE International Conference on Computer Vision (ICCV), 2019

  41. [49]

    Bottom-up object detection by grouping extreme and center points,

    X. Zhou, J. Zhuo, and P . Krahenbuhl, “Bottom-up object detection by grouping extreme and center points,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  42. [50]

    Associative embedding: End- to-end learning for joint detection and grouping,

    A. Newell, Z. Huang, and J. Deng, “Associative embedding: End- to-end learning for joint detection and grouping,” in Advances in Neural Information Processing Systems (NIPS), 2017

  43. [51]

    The pascal visual object classes (voc) challenge,

    M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” International Journal of Computer Vision (IJCV) , vol. 88, no. 2, pp. 303–338, 2010

  44. [52]

    Ima- geNet: A Large-Scale Hierarchical Image Database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Ima- geNet: A Large-Scale Hierarchical Image Database,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2009

  45. [53]

    Bounding box regression with uncertainty for accurate object detection,

    Y. He, C. Zhu, J. Wang, M. Savvides, and X. Zhang, “Bounding box regression with uncertainty for accurate object detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  46. [54]

    Is sampling heuristics necessary in training deep object detectors?

    J. Chen, D. Liu, T. Xu, S. Zhang, S. Wu, B. Luo, X. Peng, and E. Chen, “Is sampling heuristics necessary in training deep object detectors?” arXiv, vol. 1909.04868, 2019

  47. [55]

    Human face detection in visual scenes,

    H. A. Rowley, S. Baluja, and T. Kanade, “Human face detection in visual scenes,” in The Advances in Neural Information Processing Systems (NIPS), 1995

  48. [56]

    Ron: Reverse connection with objectness prior networks for object detection,

    T. Kong, F. Sun, A. Yao, H. Liu, M. Lu, and Y. Chen, “Ron: Reverse connection with objectness prior networks for object detection,” OKSUZ et al.: IMBALANCE PROBLEMS IN OBJECT DETECTION: A REVIEW 31 in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017

  49. [57]

    Single-shot refinement neural network for object detection,

    S. Zhang, L. Wen, X. Bian, Z. Lei, and S. Z. Li, “Single-shot refinement neural network for object detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018

  50. [58]

    Enriched feature guided refinement network for object detection,

    J. Nie, R. M. Anwer, H. Cholakkal, F. S. Khan, Y. Pang, and L. Shao, “Enriched feature guided refinement network for object detection,” in The IEEE International Conference on Computer Vision (ICCV), 2019

  51. [59]

    Gradient harmonized single-stage detector,

    B. Li, Y. Liu, and X. Wang, “Gradient harmonized single-stage detector,” in AAAI Conference on Artificial Intelligence , 2019

  52. [60]

    Residual objectness for imbalance reduction,

    J. Chen, D. Liu, B. Luo, X. Peng, T. Xu, and E. Chen, “Residual objectness for imbalance reduction,” arXiv, vol. 1908.09075, 2019

  53. [61]

    Towards accurate one-stage object detection with ap-loss,

    K. Chen, J. Li, W. Lin, J. See, J. Wang, L. Duan, Z. Chen, C. He, and J. Zou, “Towards accurate one-stage object detection with ap-loss,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  54. [62]

    DR Loss: Improving Object Detection by Distributional Ranking,

    Q. Qian, L. Chen, H. Li, and R. Jin, “DR Loss: Improving Object Detection by Distributional Ranking,” arXiv, vol. 1907.10156, 2019

  55. [63]

    A-fast-rcnn: Hard pos- itive generation via adversary for object detection,

    X. Wang, A. Shrivastava, and A. Gupta, “A-fast-rcnn: Hard pos- itive generation via adversary for object detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017

  56. [64]

    Learning to generate synthetic data via compositing,

    S. Tripathi, S. Chandra, A. Agrawal, A. Tyagi, J. M. Rehg, and V . Chari, “Learning to generate synthetic data via compositing,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  57. [65]

    Data augmentation for object detection via progressive and selective instance-switching,

    H. Wang, Q. Wang, F. Yang, W. Zhang, and W. Zuo, “Data augmentation for object detection via progressive and selective instance-switching,” arXiv, vol. 1906.00358, 2019

  58. [66]

    Generating posi- tive bounding boxes for balanced training of object detectors,

    K. Oksuz, B. C. Cam, S. Kalkan, and E. Akbas, “Generating posi- tive bounding boxes for balanced training of object detectors,” in IEEE Winter Applications on Computer Vision (WACV), 2020

  59. [67]

    Region proposal by guided anchoring,

    J. Wang, K. Chen, S. Yang, C. C. Loy, and D. Lin, “Region proposal by guided anchoring,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  60. [68]

    Freeanchor: Learning to match anchors for visual object detection,

    X. Zhang, F. Wan, C. Liu, R. Ji, and Q. Ye, “Freeanchor: Learning to match anchors for visual object detection,” in Advances in Neural Information Processing Systems (NIPS), 2019

  61. [69]

    Exploit all the layers: Fast and accurate cnn object detector with scale dependent pooling and cascaded rejection classifiers,

    F. Yang, W. Choi, and Y. Lin, “Exploit all the layers: Fast and accurate cnn object detector with scale dependent pooling and cascaded rejection classifiers,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016

  62. [70]

    A unified multi- scale deep convolutional neural network for fast object detec- tion,

    Z. Cai, Q. Fan, R. Feris, and N. Vasconcelos, “A unified multi- scale deep convolutional neural network for fast object detec- tion,” in The European Conference on Computer Vision (ECCV), 2016

  63. [71]

    Scale-aware fast r-cnn for pedestrian detection,

    J. Li, X. Liang, S. Shen, T. Xu, J. Feng, and S. Yan, “Scale-aware fast r-cnn for pedestrian detection,” IEEE Transactions on Multimedia , vol. 20, no. 4, pp. 985–996, 2018

  64. [72]

    Efficient featurized image pyramid network for single shot detector,

    Y. Pang, T. Wang, R. M. Anwer, F. S. Khan, and L. Shao, “Efficient featurized image pyramid network for single shot detector,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  65. [73]

    Better to follow, follow to be better: Towards precise supervision of feature super- resolution for small object detection,

    J. Noh, W. Bae, W. Lee, J. Seo, and G. Kim, “Better to follow, follow to be better: Towards precise supervision of feature super- resolution for small object detection,” in The IEEE International Conference on Computer Vision (ICCV), 2019

  66. [74]

    Scale-aware trident net- works for object detection,

    Y. Li, Y. Chen, N. Wang, and Z. Zhang, “Scale-aware trident net- works for object detection,” in The IEEE International Conference on Computer Vision (ICCV), 2019

  67. [75]

    Path aggregation network for instance segmentation,

    S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path aggregation network for instance segmentation,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  68. [76]

    Scale-transferrable object detection,

    P . Zhou, B. Ni, C. Geng, J. Hu, and Y. Xu, “Scale-transferrable object detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  69. [77]

    Parallel feature pyramid network for object detection,

    S.-W. Kim, H.-K. Kook, J.-Y. Sun, M.-C. Kang, and S.-J. Ko, “Parallel feature pyramid network for object detection,” in The European Conference on Computer Vision (ECCV), 2018

  70. [78]

    Deep feature pyramid reconfiguration for object detection,

    T. Kong, F. Sun, W. Huang, and H. Liu, “Deep feature pyramid reconfiguration for object detection,” in The European Conference on Computer Vision (ECCV), 2018

  71. [79]

    Zoom out-and-in network with map attention decision for region proposal and object detection,

    H. Li, Y. Liu, W. Ouyang, and X. Wang, “Zoom out-and-in network with map attention decision for region proposal and object detection,” International Journal of Computer Vision, vol. 127, no. 3, pp. 225–238, 2019

  72. [80]

    M2det: A single-shot object detector based on multi-level feature pyramid network,

    Q. Zhao, T. Sheng, Y. Wang, Z. Tang, Y. Chen, L. Cai, and H. Ling, “M2det: A single-shot object detector based on multi-level feature pyramid network,” in AAAI Conference on Artificial Intelligence , 2019

  73. [81]

    NAS-FPN: learning scalable feature pyramid architecture for object detection,

    G. Ghiasi, T. Lin, R. Pang, and Q. V . Le, “NAS-FPN: learning scalable feature pyramid architecture for object detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  74. [82]

    Auto-fpn: Automatic network architecture adaptation for object detection beyond classification,

    H. Xu, L. Yao, W. Zhang, X. Liang, and Z. Li, “Auto-fpn: Automatic network architecture adaptation for object detection beyond classification,” in The IEEE International Conference on Computer Vision (ICCV), 2019

  75. [83]

    Unitbox: An advanced object detection network,

    J. Yu, Y. Jiang, Z. Wang, Z. Cao, and T. Huang, “Unitbox: An advanced object detection network,” in The ACM International Conference on Multimedia, 2016

  76. [84]

    Improving object local- ization with fitness nms and bounded iou loss,

    L. Tychsen-Smith and L. Petersson, “Improving object local- ization with fitness nms and bounded iou loss,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018

  77. [85]

    Generalized intersection over union: A metric and a loss for bounding box regression,

    H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese, “Generalized intersection over union: A metric and a loss for bounding box regression,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  78. [86]

    Distance-iou loss: Faster and better learning for bounding box regression,

    Z. Zheng, P . Wang, W. Liu, J. Li, Y. Rongguang, and R. Dongwei, “Distance-iou loss: Faster and better learning for bounding box regression,” in AAAI Conference on Artificial Intelligence , 2020

  79. [87]

    Cascade R-CNN: Delving into high quality object detection,

    Z. Cai and N. Vasconcelos, “Cascade R-CNN: Delving into high quality object detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  80. [88]

    Hierarchical shot detector,

    J. Cao, Y. Pang, J. Han, and X. Li, “Hierarchical shot detector,” in The IEEE International Conference on Computer Vision (ICCV), 2019

  81. [89]

    Iou-uniform r-cnn: Breaking through the limitations of rpn,

    Z. Li, Z. Xie, L. Liu, B. Tao, and W. Tao, “Iou-uniform r-cnn: Breaking through the limitations of rpn,” arXiv, vol. 1912.05190, 2019

  82. [90]

    Microsoft COCO: Common Ob- jects in Context,

    T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P . Perona, D. Ramanan, P . Dollár, and C. L. Zitnick, “Microsoft COCO: Common Ob- jects in Context,” in The European Conference on Computer Vision (ECCV), 2014

  83. [91]

    The open images dataset V4: unified image classifica- tion, object detection, and visual relationship detection at scale,

    A. Kuznetsova, H. Rom, N. Alldrin, J. R. R. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, T. Duerig, and V . Ferrari, “The open images dataset V4: unified image classifica- tion, object detection, and visual relationship detection at scale,” arXiv, vol. 18...

  84. [92]

    Objects365: A large-scale, high-quality dataset for object detection,

    S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun, “Objects365: A large-scale, high-quality dataset for object detection,” in The IEEE International Conference on Computer Vision (ICCV), 2019

  85. [93]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016

  86. [94]

    Example-based learning for view- based human face detection,

    K.-K. Sung and T. Poggio, “Example-based learning for view- based human face detection,” IEEE Transactions on Pattern Anal- ysis and Machine Intelligence (TP AMI) , vol. 20, no. 1, pp. 39–51, 1998

  87. [95]

    Rapid object detection using a boosted cascade of simple features,

    P . Viola and M. Jones, “Rapid object detection using a boosted cascade of simple features,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2001

  88. [96]

    Histograms of oriented gradients for human detection,

    N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2005

  89. [97]

    Training deep neural networks via direct loss minimization,

    Y. Song, A. Schwing, Richard, and R. Urtasun, “Training deep neural networks via direct loss minimization,” inThe International Conference on Machine Learning (ICML), 2016

  90. [98]

    End-to-end training of object class detectors for mean average precision,

    P . Henderson and V . Ferrari, “End-to-end training of object class detectors for mean average precision,” in The Asian Conference on Computer Vision (ACCV), 2017

  91. [99]

    Cut, paste and learn: Surprisingly easy synthesis for instance detection,

    D. Dwibedi, I. Misra, and M. Hebert, “Cut, paste and learn: Surprisingly easy synthesis for instance detection,” in The IEEE International Conference on Computer Vision (ICCV), 2017

  92. [100]

    Modeling visual context is key to augmenting object detection datasets,

    N. Dvornik, J. Mairal, and C. Schmid, “Modeling visual context is key to augmenting object detection datasets,” in The European Conference on Computer Vision (ECCV), 2018

  93. [101]

    Going deeper with OKSUZ et al.: IMBALANCE PROBLEMS IN OBJECT DETECTION: A REVIEW 32 convolutions,

    C. Szegedy, W. Liu, Y. Jia, P . Sermanet, S. Reed, D. Anguelov, D. Erhan, V . Vanhoucke, and A. Rabinovich, “Going deeper with OKSUZ et al.: IMBALANCE PROBLEMS IN OBJECT DETECTION: A REVIEW 32 convolutions,” in The IEEE Conference on Computer Vision and Pattern Recognition (CV...

  94. [102]

    Localization recall precision (LRP): A new performance metric for object detection,

    K. Oksuz, B. C. Cam, E. Akbas, and S. Kalkan, “Localization recall precision (LRP): A new performance metric for object detection,” in The European Conference on Computer Vision (ECCV), 2018

  95. [103]

    Very deep convolutional net- works for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional net- works for large-scale image recognition,” in The International Conference on Learning Representations (ICLR), 2015

  96. [104]

    Stacked hourglass networks for human pose estimation,

    A. Newell, K. Yang, and J. Deng, “Stacked hourglass networks for human pose estimation,” in The European Conference on Computer Vision (ECCV), B. Leibe, J. Matas, N. Sebe, and M. Welling, Eds., 2016

  97. [105]

    Mobilenets: Efficient convolutional neural networks for mobile vision applications,

    A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017

  98. [106]

    Detnet: Design backbone for object detection,

    Z. Li, C. Peng, G. Yu, X. Zhang, Y. Deng, and J. Sun, “Detnet: Design backbone for object detection,” in The European Conference on Computer Vision (ECCV), 2018

  99. [107]

    Bidirectional feature pyramid network with recurrent attention residual modules for shadow detection,

    L. Zhu, Z. Deng, X. Hu, C.-W. Fu, X. Xu, J. Qin, and P .-A. Heng, “Bidirectional feature pyramid network with recurrent attention residual modules for shadow detection,” in The European Confer- ence on Computer Vision (ECCV), 2018

  100. [108]

    Concatenated feature pyramid network for instance segmentation,

    Y. Sun, P . S. K. P , J. Shimamura, and A. Sagata, “Concatenated feature pyramid network for instance segmentation,” arXiv, vol. 1904.00768, 2019

  101. [109]

    Feature pyramid network for multi-class land segmentation,

    S. S. Seferbekov, V . I. Iglovikov, A. V . Buslaev, and A. A. Shvets, “Feature pyramid network for multi-class land segmentation,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2018

  102. [110]

    Panoptic feature pyramid networks,

    A. Kirillov, R. B. Girshick, K. He, and P . Dollár, “Panoptic feature pyramid networks,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  103. [111]

    Pyramid methods in image processing,

    E. H. Adelson, C. H. Anderson, J. R. Bergen, P . J. Burt, and J. M. Ogden, “Pyramid methods in image processing,” RCA Engineer, vol. 29, no. 6, pp. 33–41, 1984

  104. [112]

    Multi-scale context aggregation by dilated convolutions,

    F. Yu and V . Koltun, “Multi-scale context aggregation by dilated convolutions,” in The International Conference on Learning Repre- sentations (ICLR), 2016

  105. [113]

    Non-local neural networks,

    X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  106. [114]

    Densely connected convolutional networks,

    G. Huang, Z. Liu, and K. Q. Weinberger, “Densely connected convolutional networks,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017

  107. [115]

    Spatial pyramid pooling in deep convolutional networks for visual recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” in The European Conference on Computer Vision (ECCV), D. Fleet, T. Pajdla, B. Schiele, and T. Tuytelaars, Eds., 2014

  108. [116]

    Squeeze-and-excitation networks,

    J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  109. [117]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift,

    S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in The International Conference on Machine Learning (ICML) , 2015

  110. [118]

    Neural architecture search with rein- forcement learning,

    B. Zoph and Q. V . Le, “Neural architecture search with rein- forcement learning,” in The International Conference on Learning Representations (ICLR), 2017

  111. [119]

    Learning trans- ferable architectures for scalable image recognition,

    B. Zoph, V . Vasudevan, J. Shlens, and Q. V . Le, “Learning trans- ferable architectures for scalable image recognition,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018

  112. [120]

    Regularized evolution for image classifier architecture search,

    E. Real, A. Aggarwal, Y. Huang, and Q. V . Le, “Regularized evolution for image classifier architecture search,” in AAAI Con- ference on Artificial Intelligence, 2019

  113. [121]

    Efficientnet: Rethinking model scaling for convolutional neural networks,

    M. Tan and Q. V . Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in The International Conference on Machine Learning (ICML), 2019

  114. [122]

    Detnas: Backbone search for object detection,

    Y. Chen, T. Yang, X. Zhang, G. Meng, C. Pan, and J. Sun, “Detnas: Backbone search for object detection,” in Advances in Neural Information Processing Systems (NIPS), 2019

  115. [123]

    Overfeat: Integrated recognition, localization and de- tection using convolutional networks,

    P . Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. Lecun, “Overfeat: Integrated recognition, localization and de- tection using convolutional networks,” in International Conference on Learning Representations (ICLR), 2014

  116. [124]

    Robust estimation of a location parameter,

    P . J. Huber, “Robust estimation of a location parameter,” Annals of Statistics, vol. 53, no. 1, pp. 73–101, 1964

  117. [125]

    Acquisition of localization confidence for accurate object detection,

    B. Jiang, R. Luo, J. Mao, T. Xiao, and Y. Jiang, “Acquisition of localization confidence for accurate object detection,” in The European Conference on Computer Vision (ECCV), 2018

  118. [126]

    Precise detection in densely packed scenes,

    E. Goldman, R. Herzig, A. Eisenschtat, J. Goldberger, and T. Has- sner, “Precise detection in densely packed scenes,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019

  119. [127]

    Gaussian yolov3: An accurate and fast object detector using localization uncertainty for autonomous driving,

    J. Choi, D. Chun, H. Kim, and H.-J. Lee, “Gaussian yolov3: An accurate and fast object detector using localization uncertainty for autonomous driving,” in The IEEE International Conference on Computer Vision (ICCV), 2019

  120. [128]

    Learning to rank pro- posals for object detection,

    Z. Tan, X. Nie, Q. Qian, N. Li, and H. Li, “Learning to rank pro- posals for object detection,” in The IEEE International Conference on Computer Vision (ICCV), 2019

  121. [129]

    Object detection via a multi- region and semantic segmentation-aware cnn model,

    S. Gidaris and N. Komodakis, “Object detection via a multi- region and semantic segmentation-aware cnn model,” in The IEEE International Conference on Computer Vision (ICCV), 2015

  122. [130]

    Attend refine repeat: Active box proposal generation via in-out localization,

    ——, “Attend refine repeat: Active box proposal generation via in-out localization,” in The British Machine Vision Conference (BMVC), 2016

  123. [131]

    Deformable convolutional networks,

    J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” in The IEEE International Conference on Computer Vision (ICCV), 2017

  124. [132]

    MMDetection: Open mmlab detection toolbox and benchmark,

    K. Chen, J. Wang, J. Pang, Y. Cao, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Xu, Z. Zhang, D. Cheng, C. Zhu, T. Cheng, Q. Zhao, B. Li, X. Lu, R. Zhu, Y. Wu, J. Dai, J. Wang, J. Shi, W. Ouyang, C. C. Loy, and D. Lin, “MMDetection: Open mmlab detection toolbox and benchmark,”...

  125. [133]

    Metaanchor: Learning to detect objects with customized anchors,

    T. Yang, X. Zhang, Z. Li, W. Zhang, and J. Sun, “Metaanchor: Learning to detect objects with customized anchors,” in Advances in Neural Information Processing Systems (NIPS) , 2018

  126. [134]

    Megdet: A large mini-batch object detector,

    C. Peng, T. Xiao, Z. Li, Y. Jiang, X. Zhang, K. Jia, G. Yu, and J. Sun, “Megdet: A large mini-batch object detector,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018

  127. [135]

    Dynamic task prioritization for dynamic task learning,

    M. Guo, A. Haque, D.-A. Huang, S. Yeung, and L. Fei-Fei, “Dynamic task prioritization for dynamic task learning,” in The European Conference on Computer Vision (ECCV), 2018

  128. [136]

    Smote: Synthetic minority over-sampling technique,

    N. V . Chawla, K. W. Bowyer, L. O. Hall, and W. P . Kegelmeyer, “Smote: Synthetic minority over-sampling technique,” Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, 2002

  129. [137]

    C4.5, class imbalance, and cost sensitivity: Why under-sampling beats over-sampling,

    C. Drummond and R. C. Holte, “C4.5, class imbalance, and cost sensitivity: Why under-sampling beats over-sampling,” in The International Conference on Machine Learning (ICML) , 2003

  130. [138]

    Adasyn: Adaptive synthetic sampling approach for imbalanced learning,

    H. He, Y. Bai, E. A. Garcia, and S. Li, “Adasyn: Adaptive synthetic sampling approach for imbalanced learning,” inInternational Joint Conference on Neural Networks, 2008

  131. [139]

    A systematic study of the class imbalance problem in convolutional neural networks,

    M. Buda, A. Maki, and M. A. Mazurowski, “A systematic study of the class imbalance problem in convolutional neural networks,” Neural Networks, vol. 106, pp. 249–259, 2018

  132. [140]

    Relay Backpropagation for Effective Learning of Deep Convolutional Neural Networks,

    S. Li, L. Zhouche, and H. Qingming, “Relay Backpropagation for Effective Learning of Deep Convolutional Neural Networks,” in The European Conference on Computer Vision (ECCV), 2016

  133. [141]

    Learning to model the tail,

    Y.-X. Wang, D. Ramanan, and M. Hebert, “Learning to model the tail,” in Advances in Neural Information Processing Systems (NIPS) , 2017

  134. [142]

    Large scale fine-grained categorization and domain-specific transfer learning,

    Y. Cui, Y. Song, C. Sun, A. G. Howard, and S. J. Belongie, “Large scale fine-grained categorization and domain-specific transfer learning,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  135. [143]

    Learning deep repre- sentation for imbalanced classification,

    C. Huang, Y. Li, C. C. Loy, and X. Tang, “Learning deep repre- sentation for imbalanced classification,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016

  136. [144]

    Class-balanced loss based on effective number of samples,

    Y. Cui, M. Jia, T. Lin, Y. Song, and S. J. Belongie, “Class-balanced loss based on effective number of samples,” in The IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) , 2019

  137. [145]

    Self-paced learning for latent variable models,

    M. P . Kumar, B. Packer, and D. Koller, “Self-paced learning for latent variable models,” in Advances in Neural Information Processing Systems (NIPS), 2010

  138. [146]

    Hide-and-seek: Forcing a net- work to be meticulous for weakly-supervised object and action localization,

    K. Kumar Singh and Y. Jae Lee, “Hide-and-seek: Forcing a net- work to be meticulous for weakly-supervised object and action localization,” in The IEEE International Conference on Computer Vision (ICCV), 2017. OKSUZ et al.: IMBALANCE PROBLEMS IN OBJECT DETECTION: A REVIEW 33

  139. [147]

    Class rectification hard mining for imbalanced deep learning,

    Q. Dong, S. Gong, and X. Zhu, “Class rectification hard mining for imbalanced deep learning,” in The International Conference on Computer Vision (ICCV), 2017

  140. [148]

    Active bias: Training more accurate neural networks by emphasizing high variance samples,

    H.-S. Chang, E. Learned-Miller, and A. McCallum, “Active bias: Training more accurate neural networks by emphasizing high variance samples,” in Advances in Neural Information Processing Systems (NIPS), 2017

  141. [149]

    Submodularity in data subset selection and active learning,

    K. Wei, R. Iyer, and J. Bilmes, “Submodularity in data subset selection and active learning,” in The International Conference on Machine Learning (ICML), 2015, pp. 1954–1963

  142. [150]

    Learning from less data: Diversified subset selection and active learning in image classification tasks,

    V . Kaushal, A. Sahoo, K. Doctor, N. R. Uppalapati, S. Shetty, P . Singh, R. K. Iyer, and G. Ramakrishnan, “Learning from less data: Diversified subset selection and active learning in image classification tasks,” arXiv, vol. 1805.11191, 2018

  143. [151]

    Active learning for convolutional neu- ral networks: A core-set approach,

    O. Sener and S. Savarese, “Active learning for convolutional neu- ral networks: A core-set approach,” in The International Conference on Learning Representations (ICLR), 2018

  144. [152]

    Are all training examples created equal? an empirical study,

    K. Vodrahalli, K. Li, and J. Malik, “Are all training examples created equal? an empirical study,” arXiv, vol. 1811.12569, 2018

  145. [153]

    Semantic redundancies in image-classification datasets: The 10% you don’t need,

    V . Birodkar, H. Mobahi, and S. Bengio, “Semantic redundancies in image-classification datasets: The 10% you don’t need,” arXiv, vol. 1901.11409, 2019

  146. [154]

    Exploring the Limits of Weakly Supervised Pretraining,

    D. Mahajan, R. Girshick, V . Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten, “Exploring the Limits of Weakly Supervised Pretraining,” in The European Conference on Computer Vision (ECCV), 2018

  147. [155]

    Selectnet: Learning to sample from the wild for imbalanced data training,

    Y. Liu, T. Gao, and H. Yang, “Selectnet: Learning to sample from the wild for imbalanced data training,” arXiv, vol. 1905.09872, 2019

  148. [156]

    Generative Adversarial Minority Oversampling,

    S. S. Mullick, S. Datta, and S. Das, “Generative Adversarial Minority Oversampling,” arXiv, vol. 1903.09730, 2019

  149. [157]

    Data augmentation in emo- tion classification using generative adversarial networks,

    X. Zhu, Y. Liu, Z. Qin, and J. Li, “Data augmentation in emo- tion classification using generative adversarial networks,” arXiv preprint arXiv:1711.00648, 2017

  150. [158]

    Gan augmentation: augmenting training data using genera- tive adversarial networks,

    C. Bowles, L. Chen, R. Guerrero, P . Bentley, R. Gunn, A. Ham- mers, D. A. Dickie, M. V . Hernández, J. Wardlaw, and D. Rueck- ert, “Gan augmentation: augmenting training data using genera- tive adversarial networks,” arXiv preprint arXiv:1810.10863, 2018

  151. [159]

    Range loss for deep face recognition with long-tailed training data,

    X. Zhang, Z. Fang, Y. Wen, Z. Li, and Y. Qiao, “Range loss for deep face recognition with long-tailed training data,” in The IEEE International Conference on Computer Vision (ICCV), 2016

  152. [160]

    Feature transfer learning for deep face recognition with long-tail data,

    X. Yin, X. Yu, K. Sohn, X. Liu, and M. K. Chandraker, “Feature transfer learning for deep face recognition with long-tail data,” arXiv, vol. 1803.09014, 2018

  153. [161]

    Deep imbalanced learning for face recognition and attribute prediction,

    C. Huang, Y. Li, C. L. Chen, and X. Tang, “Deep imbalanced learning for face recognition and attribute prediction,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TP AMI), 2019

  154. [162]

    Discriminative deep metric learning for face verification in the wild,

    J. Hu, J. Lu, and Y. Tan, “Discriminative deep metric learning for face verification in the wild,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014

  155. [163]

    Learning visual similarity for product design with convolutional neural networks,

    S. Bell and K. Bala, “Learning visual similarity for product design with convolutional neural networks,” ACM Trans. on Graphics (SIGGRAPH), vol. 34, no. 4, 2015

  156. [164]

    Facenet: A unified embedding for face recognition and clustering,

    F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in Conference on Computer Vision and Pattern Recognition (CVPR), 2015

  157. [165]

    Deep metric learning using triplet net- work,

    E. Hoffer and N. Ailon, “Deep metric learning using triplet net- work,” in The International Conference on Learning Representations (ICLR), 2015

  158. [166]

    Deep metric learning with hierarchical triplet loss,

    W. Ge, “Deep metric learning with hierarchical triplet loss,” in The European Conference on Computer Vision (ECCV), 2018

  159. [167]

    Hard-aware deeply cascaded embedding,

    Y. Yuan, K. Yang, and C. Zhang, “Hard-aware deeply cascaded embedding,” in The IEEE International Conference on Computer Vision (ICCV), 2017

  160. [168]

    Smart mining for deep metric learning,

    B. Harwood, G. K. B. VijayKumarB., G. Carneiro, I. D. Reid, and T. Drummond, “Smart mining for deep metric learning,” in The IEEE International Conference on Computer Vision (ICCV), 2017

  161. [169]

    Fine-grained categoriza- tion and dataset bootstrapping using deep metric learning with humans in the loop,

    Y. Cui, F. Zhou, Y. Lin, and S. Belongie, “Fine-grained categoriza- tion and dataset bootstrapping using deep metric learning with humans in the loop,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016

  162. [170]

    Deep metric learning via lifted structured feature embedding,

    H. O. Song, Y. Xiang, S. Jegelka, and S. Savarese, “Deep metric learning via lifted structured feature embedding,” in The IEEE Computer Vision and Pattern Recognition (CVPR), 2016

  163. [171]

    Local similarity-aware deep feature embedding,

    C. Huang, C. C. Loy, and X. Tang, “Local similarity-aware deep feature embedding,” in Advances in Neural Information Processing Systems (NIPS), 2016

  164. [172]

    Deep adversarial metric learning,

    Y. Duan, W. Zheng, X. Lin, J. Lu, and J. Zhou, “Deep adversarial metric learning,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  165. [173]

    An adversarial approach to hard triplet generation,

    Y. Zhao, Z. Jin, G.-j. Qi, H. Lu, and X.-s. Hua, “An adversarial approach to hard triplet generation,” in The European Conference on Computer Vision (ECCV), 2018

  166. [174]

    Generative adver- sarial nets,

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde- Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adver- sarial nets,” in Advances in Neural Information Processing Systems (NIPS), 2014

  167. [175]

    Hardness-aware deep metric learning,

    W. Zheng, Z. Chen, J. Lu, and J. Zhou, “Hardness-aware deep metric learning,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  168. [176]

    Interactive object detection,

    A. Yao, “Interactive object detection,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012

  169. [177]

    Self-paced multi-task learning,

    C. Li, J. Yan, F. Wei, W. Dong, Q. Liu, and H. Zha, “Self-paced multi-task learning,” in AAAI Conference on Artificial Intelligence , 2017

  170. [178]

    Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,

    A. Kendall, Y. Gal, and R. Cipolla, “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  171. [179]

    End-to-end multi-task learn- ing with attention,

    S. Liu, E. Johns, and A. J. Davison, “End-to-end multi-task learn- ing with attention,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  172. [180]

    Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks,

    C.-Y. L. Zhao Chen, Vijay Badrinarayanan and A. Rabinovich, “Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks,” in International Conference on Ma- chine Learning (ICML), 2018. Kemal Oksuz received B.Sc. in System Engi- neering from Land F...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.