Pith. sign in

REVIEW 5 major objections 5 minor 278 references

Object Recognition Datasets and Challenges: A Review

T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This survey maps more than 160 public object recognition datasets, the major benchmarking competitions, and the evaluation metrics that make results comparable.

desk verdict Broad survey with a useful index, but the table statistics are too error-prone to rely on in the current version. read the letter →

arxiv 2507.22361 v1 pith:IROE4WC4 submitted 2025-07-30 cs.CV cs.AI

classification cs.CVcs.AI
keywords objectrecognitiondatasetsurveybenchmarksevaluationmetricsdeeplearningcomputervisionautonomousdrivingdatasetsface
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to be a reference catalog of object recognition datasets, and it argues that the field's progress is best understood through the data rather than only algorithms. The authors review more than 160 annotated 2D image datasets in the visible spectrum, organizing them into generic recognition, scene understanding, and six application areas: autonomous driving, medical imaging, face recognition, remote sensing, species recognition, and clothing detection. They also describe the four major large-scale benchmarks (PASCAL VOC, ImageNet, MS COCO, Open Images), their annual challenges, and the evaluation metrics that make results comparable. A reader who wants to choose a dataset or understand why benchmarks keep changing would use this as a map.

What carries the argument

The carrying objects are the comparative dataset tables and the task-and-metric taxonomy. Table 1 anchors the four large-scale generic datasets; later tables organize detection, segmentation, scene parsing, salient-object, and application-specific datasets; and the metric section defines AP, mAP, IoU, Panoptic Quality, and OKS, linking each to the challenge that uses it. The taxonomy itself—generic versus fine-grained, bounding box versus mask, classification versus detection versus segmentation—is what turns scattered dataset papers into a navigable reference.

What would settle it

Pick 30 datasets at random from the tables, find their original publications, and compare the paper's stated image count, class count, and annotation count; if even a handful disagree, the survey's promise as a reliable quantitative reference fails, because the tables carry the argument and most rows cite no source.

Watch

Extended reading notes

Core claim

The central claim is that the dataset landscape, though vast, can be systematically catalogued and compared. The paper presents statistics on image counts, class counts, annotation counts, and annotation types for more than 160 public datasets, identifies the dominant trends (from iconic single-object classification to cluttered, pixel-level, context-rich annotation), and documents the competition tasks and metric formulas that define state-of-the-art evaluation. Read fairly, the paper says that datasets are the load-bearing infrastructure of object recognition progress: as algorithms saturate existing benchmarks, the community responds by building larger, finer-grained, and more contextually rich datasets, and this cyclic process is visible in the tables and challenge histories.

Load-bearing premise

The tables' statistics on image counts, class counts, and annotation counts were transcribed accurately from the original dataset publications, even though most rows cite no source.

Editorial extensions

If this is right

  • Practitioners can use the survey as a first screening step to narrow candidate datasets by task, size, annotation type, and application.
  • The historical trajectory shows annotation moving from image-level labels toward pixel-level and instance-aware masks, so newer benchmarks should be expected to follow that pattern.
  • Saturation of existing benchmarks creates pressure for new, harder datasets, as the paper argues.
  • Standardized metrics such as mAP with IoU thresholds, Panoptic Quality, and OKS make challenge results comparable across years.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The survey's coverage is citation-weighted, so lesser-known but technically valuable datasets may be under-represented; a systematic search protocol would make future updates more reproducible.
  • The paper's tables could be turned into a queryable database, letting readers filter by class count, annotation type, and application, which would extend its usefulness beyond the printed page.
  • If the field keeps producing datasets faster than surveys, a living document updated on release cadence would serve the community better than a one-time review.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. This manuscript is a survey-style reference catalog of object recognition datasets, benchmarks, and evaluation metrics. It covers generic large-scale datasets (PASCAL VOC, ImageNet, COCO, Open Images), additional detection/segmentation/scene datasets, and six application areas (autonomous driving, medical imaging, face recognition, remote sensing, species recognition, clothing detection), with tabulated statistics and challenge descriptions, plus a GitHub repository. The stated contribution is a quantitative and descriptive analysis of more than 160 datasets.

Significance. If accurate, the paper would be a useful starting point for practitioners selecting benchmarks and for researchers seeking an overview of dataset trends, because it consolidates statistics, challenge tasks, and metric formulas across many domains in one place and makes the catalog available online. The breadth is a genuine strength: the tables cover well-known and less standard datasets, and the metric section collects formulas (precision, recall, AP, mAP, OKS, PQ, etc.) in a single reference. However, the paper's value rests entirely on the accuracy of the compiled statistics and citations; the factual errors identified below (COCO class counts, ImageNet object counts and year, DRIVE size/year, Eq. (13) labeling, SSD attribution) mean the catalog is not yet reliable in its current form.

major comments (5)
  1. [Section 3.1, Table 1; Section 3.1.2, Table 2; Section 3.3, Table 6] The class count for Microsoft COCO is inconsistent and incorrect in the headline table. Table 1 reports 91 classes, but the COCO object detection benchmark uses 80 thing classes; Table 2's Detection row correctly lists 80, and the Stuff row lists 91 stuff classes. Section 3.3's COCO Stuff row also separates 91 stuff from 80 thing classes. Table 1 should either state 80 object classes or explicitly distinguish the 80 thing classes from the 91 stuff classes.
  2. [Section 3.1.1 and Table 1] ImageNet statistics are internally inconsistent. Section 3.1.1 says ImageNet suffers from 'partial image annotations (only one annotated object per image)', while Table 1 reports an average of 3 objects per image. The same paragraph says ImageNet was introduced in 2010, while Table 1 lists 2009; the text appears to conflate the dataset with the ILSVRC challenge. These numbers must be reconciled and sourced.
  3. [Section 4.2, Tables 12 and 13] The DRIVE entry is factually wrong: Table 12 lists '400 cases' with year 2019 and describes images of 400 patients, and Table 13 repeats the 2019 date. The DRIVE vessel-extraction database contains 40 retinal images and was published in the early 2000s (the cited reference [219] is the DRIVE paper). Since the paper's contribution is a reliable quantitative catalog, every medical imaging row needs verification against its source; this error strongly suggests that unverified rows may contain similar mistakes.
  4. [Section 3.1.2, Eq. (13)] The equation following the top-5 error description is labeled 'P Q', but it is not the Panoptic Quality metric. It appears to be a top-k classification error formula (mean over test examples of the minimum 0/1 distance between predicted and ground-truth labels). This mislabeling will confuse readers; the equation should be renamed 'top-5 error' and the variables (l_j, g_k, n, k) defined precisely.
  5. [General, Tables 1-18] Most dataset tables lack per-row citations. Tables 1, 3-10, 12, 14, 15, 17, and 18 list sizes, class counts, and years without a source column, and some rows have no associated reference anywhere in the bibliography. For a reference catalog, this makes it impossible for readers to distinguish isolated typos from systematic transcription errors. I request row-level citations, or a supplementary source table, for every numerical entry.
minor comments (5)
  1. [Section 2.2] SSD is attributed to 'C. Szegedy et al.' with reference [158], but the cited work is Liu et al., 'SSD: Single Shot MultiBox Detector'; the attribution and reference should be corrected.
  2. [Section 2.3, Eq. (11)] In the definition of Panoptic Quality, the text describes FP as 'incorrect detections with IoU > 0.5'; a false positive should be a detection with IoU at or below the 0.5 threshold, so this is likely a typo that should be fixed.
  3. [Section 1 and figures] The figure references are off: the text cites 'Fig. 2' for the ROC curve (the caption says Figure 3), and Section 3.1 refers to 'Figure 3' for the challenge accuracy plot, which is captioned as Figure 4. The figures should be renumbered or the references corrected.
  4. [Throughout] There are numerous typos and minor errors, including 'withing' in Section 1, 'CIF AR-10'/'CIF AR-100' spacing in Section 2.2, 'Rential' in Table 12, 'Automative RADAR' and 'Paolo Alto' in Table 8, and 'section 3.3.1' in Section 3.1.2, which should be Section 3.1. A careful proofreading pass is needed.
  5. [References] Several references are incomplete or lack publication years (e.g., [60], [67], [239], [240]), and reference [1] is a bare URL with no title; the bibliography should be completed before publication.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey compiles external dataset statistics and contains no derivation whose output is defined by its own inputs.

full rationale

This paper is a survey/catalog rather than a derivation chain. Its central claim — that more than 160 datasets are "scrutinized through statistics and descriptions" — is supported by tables and narrative descriptions drawn from the original dataset publications; no parameter is fitted, no prediction is generated from the paper's own assumptions, and no result is defined in terms of the paper's own output. The only author self-citation in the argument, reference [63] (co-author Djavadifar's PhD thesis), supports the peripheral aside that bounding box annotations are "more robust to inconsistencies arising from subjective annotations of different human annotators working on the same dataset"; this claim is not load-bearing for any benchmark comparison or for the survey's catalog. The manuscript's demonstrable transcription errors, such as the COCO class count in Table 1 and the DRIVE size/year in Table 12, are accuracy and correctness concerns, not circularity: the statistics are intended to reproduce external facts, and a wrong reproduction is not the same as making the conclusion equivalent to the input. The survey is self-contained against external dataset publications and exhibits no reduction of a claimed result to its own inputs, so the circularity score is minimal.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The survey's conclusions rely on the accuracy of the statistics transcribed from original dataset papers and on the representativeness of the chosen dataset selection. No free parameters or invented entities are introduced.

assumptions (2)
  • domain assumption The surveyed dataset statistics accurately reflect the original dataset publications.
    The paper transcribes numbers into tables (e.g., Table 1, 8, 12) without independent verification or per-entry citations.
  • domain assumption The selection of more than 160 datasets is representative and unbiased.
    Section 1.2 restricts scope to 2D visible-spectrum annotated datasets and acknowledges omissions due to limited space; generalizability of conclusions depends on this selection.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Object Recognition Datasets and Challenges: A Review." pith.science (2026). https://pith.science/paper/IROE4WC4

@misc{pith2026250722361,
  author       = {Pith},
  title        = {Pith review of: Object Recognition Datasets and Challenges: A Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IROE4WC4}},
  note         = {Machine review of arXiv:2507.22361}
}
read the original abstract

Object recognition is among the fundamental tasks in the computer vision applications, paving the path for all other image understanding operations. In every stage of progress in object recognition research, efforts have been made to collect and annotate new datasets to match the capacity of the state-of-the-art algorithms. In recent years, the importance of the size and quality of datasets has been intensified as the utility of the emerging deep network techniques heavily relies on training data. Furthermore, datasets lay a fair benchmarking means for competitions and have proved instrumental to the advancements of object recognition research by providing quantifiable benchmarks for the developed models. Taking a closer look at the characteristics of commonly-used public datasets seems to be an important first step for data-driven and machine learning researchers. In this survey, we provide a detailed analysis of datasets in the highly investigated object recognition areas. More than 160 datasets have been scrutinized through statistics and descriptions. Additionally, we present an overview of the prominent object recognition benchmarks and competitions, along with a description of the metrics widely adopted for evaluation purposes in the computer vision community. All introduced datasets and challenges can be found online at github.com/AbtinDjavadifar/ORDC.

Figures

Figures reproduced from arXiv: 2507.22361 by the authors.

Figure 1
Figure 1. Overview of object recognition tasks in computer vision research [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Milestones in object recognition algorithm and dataset development.(MNIST [139, 1], FERET [189], COIL-20 [176], [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. ROC Curve Recall = T P T P + F N (2) Specificity (True Negative Rate): Specificity measures the share of actual negatives that have been correctly predicted as such. Specif icity = T N T N + F P (3) Accuracy: Accuracy shows the ratio of correctly predicted instances to all the available instances. Accuracy = T P + T N T P + T N + F P + F N (4) It needs to be noted that a similarity threshold between the annotated da… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Accuracy improvement of winner algorithms in the object detection track of major challenges. The fall in accuracy [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Different types of medical images included in datasets. The number of available datasets related to each organ is [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

278 extracted references · 38 canonical work pages

  1. [219]

    Locality Sensitive Deep Learning for Detection and Classification of Nuclei in Routine Colon Cancer Histology Images

    Sirinukunwattana, K., Raza, s.E.A., Tsang, Y., Snead, D.R., Cree, I.A., Rajpoot, N.M., 2016. Locality Sensitive Deep Learning for Detection and Classification of Nuclei in Routine Colon Cancer Histology Images. IEEE Transactions on Medical Imaging 35, 1196–1206. doi: 10.1109/TMI.2016.2525803

  2. [1]

    IEEE Xplore Full-Text PDF:

    , . IEEE Xplore Full-Text PDF:. URL: https://ieeexplore.ieee.org/stamp/stamp.jsp?tp={\&}arnumber=8241865

  3. [2]

    2D and 3D face recognition: A survey

    Abate, A.F., Nappi, M., Riccio, D., Sabatino, G., 2007. 2D and 3D face recognition: A survey. Pattern Recognition Letters 28, 1885–1906. doi: 10.1016/j.patrec.2006.12.018

  4. [4]

    Endoscopy artifact detection (EAD 2019) challenge dataset , 1–13doi: 10.17632/C7FJBXCGJ9.1

    Ali, S., Zhou, F., Daul, C., Braden, B., Bailey, A., Realdon, S., East, J., Wagni` eres, G., Loschenov, V., Grisan, E., Blondel, W., Rittscher, J., 2019a. Endoscopy artifact detection (EAD 2019) challenge dataset , 1–13doi: 10.17632/C7FJBXCGJ9.1

  5. [5]

    EAD 2019

    Ali, S., Zhou, F., Daul, C., Loschenov, M., 2019b. EAD 2019. URL: https://ead2019.grand-challenge.org/

  6. [6]

    Overview of artificial intelligence in Medicine

    Amisha, Malik, P., Pathania, M., Rathaur, V.K., 2019. Overview of artificial intelligence in Medicine. Journal of Family Medicine and Primary Care 8, 2328–2331. doi: 10.4103/jfmpc.jfmpc_440_19

  7. [7]

    CVPR 2019 W AD Beyond Single-frame Perception Challenge

    Apolloscape, 2019. CVPR 2019 W AD Beyond Single-frame Perception Challenge. URL: http://wad.ai/2019/index. html

  8. [8]

    ICIAR 2018

    Ara´ ujo, T., Aresta, G., Eloy, C., Ant´ onio, P., Aguiar, P., 2018. ICIAR 2018. URL: https://iciar2018-challenge. grand-challenge.org/. 22

Show all 278 references
  1. [9]

    BACH: Grand challenge on breast cancer histology images

    Aresta, G., Ara´ ujo, T., Kwok, S., Chennamsetty, S.S., Safwan, M., Alex, V., Marami, B., Prastawa, M., Chan, M., Donovan, M., Fernandez, G., Zeineh, J., Kohl, M., Walz, C., Ludwig, F., Braunewell, S., Baust, M., Vu, Q.D., To, M.N.N., Kim, E., Kwak, J.T., Galal, S., Sanchez-Fr...

  2. [10]

    The Lung Image Database Consortium (LIDC) and Image Database Resource Initiative (IDRI): A completed reference database of lung nodules on CT scans

    Armato, S.G., McLennan, G., Bidaut, L., McNitt-Gray, M.F., Meyer, C.R., Reeves, A.P., Zhao, B., Aberle, D.R., Henschke, C.I., Hoffman, E.A., Kazerooni, E.A., MacMahon, H., Van Beek, E.J., Yankelevitz, D., Biancardi, A.M., Bland, P.H., Brown, M.S., Engelmann, R.M., Laderach, G....

  3. [11]

    UMDFaces: An Annotated Face Dataset for Training Deep Networks

    Bansal, A., Nanduri, A., Castillo, C., Ranjan, R., Chellappa, R., 2016. UMDFaces: An Annotated Face Dataset for Training Deep Networks. IEEE International Joint Conference on Biometrics, IJCB 2017 2018-Janua, 464–473. URL: http://arxiv.org/abs/1611.01484

  4. [12]

    ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models

    Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., Katz, B., 2019. ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. Advances in neural information processing systems , 1–11URL: https://objectnet.dev

  5. [13]

    Surf: Speeded up robust features , 404–417

    Bay, H., Tuytelaars, T., Van Gool, L., 2006. Surf: Speeded up robust features , 404–417

  6. [14]

    SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences URL: http://arxiv.org/abs/1904.01416

    Behley, J., Garbade, M., Milioto, A., Quenzel, J., Behnke, S., Stachniss, C., Gall, J., 2019. SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences URL: http://arxiv.org/abs/1904.01416

  7. [15]

    OPENSURF ACES: A richly annotated catalog of surface appearance

    Bell, S., Upchurch, P., Snavely, N., Bala, K., 2013. OPENSURF ACES: A richly annotated catalog of surface appearance. ACM Transactions on Graphics 32. doi: 10.1145/2461912.2462002

  8. [16]

    Greedy layer-wise training of deep networks, in: Advances in neural information processing systems, pp

    Bengio, Y., Lamblin, P., Popovici, D., Larochelle, H., 2007. Greedy layer-wise training of deep networks, in: Advances in neural information processing systems, pp. 153–160

  9. [17]

    Names and faces in the news

    Berg, T.L., Berg, A.C., Edwards, J., Maire, M., White, R., Teh, Y.W., Learned-Miller, E., Forsyth, D.A., 2004. Names and faces in the news. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition

  10. [19]

    CVPR 2018 - Berkeley DeepDrive challenges

    Berkeley Deep Drive, 2018. CVPR 2018 - Berkeley DeepDrive challenges

  11. [20]

    Towards automatic polyp detection with a polyp appearance model

    Bernal, J., S´ anchez, J., Vilarino, F., 2012. Towards automatic polyp detection with a polyp appearance model. Pattern Recognition 45, 3166–3182. doi: https://doi.org/10.1016/j.patcog.2012.03.002

  12. [21]

    Automatic 3D face authentication

    Beumier, C., Acheroy, M., 2000. Automatic 3D face authentication. Image and Vision Computing 18, 315–321. doi: 10. 1016/S0262-8856(99)00052-9

  13. [22]

    Deep-learning-assisted diagnosis for knee magnetic resonance imaging: Development and retrospective validation of MRNet

    Bien, N., Rajpurkar, P., Ball, R.L., Irvin, J., Park, A., Jones, E., Bereket, M., Patel, B.N., Yeom, K.W., Shpanskaya, K., Halabi, S., Zucker, E., Fanton, G., Amanatullah, D.F., Beaulieu, C.F., Riley, G.M., Stewart, R.J., Blankenberg, F.G., Larson, D.B., Jones, R.H., Langlotz,...

  14. [23]

    The inD Dataset: A Drone Dataset of Naturalistic Road User Trajectories at German Intersections

    Bock, J., Krajewski, R., Moers, T., Runde, S., Vater, L., Eckstein, L., 2019. The inD Dataset: A Drone Dataset of Naturalistic Road User Trajectories at German Intersections

  15. [24]

    Long- term underwater camera surveillance for monitoring and analysis of fish populations

    Boom, B.J., Huang, P.X., Beyan, C., Spampinato, C., Palazzo, S., He, J., Beauxis-Aussalet, E., Lin, S.I., Chou, H.M., Nadarajan, G., Chen-Burger, Y.H., van Ossenbruggen, J., Giordano, D., Hardman, L., Lin, F.P., Fisher, R.B., 2012. Long- term underwater camera surveillance for...

  16. [25]

    Learning fuzzy concept definitions

    Botta, M., Giordana, A., Saitta, L., 1993. Learning fuzzy concept definitions. 1993 IEEE International Conference on Fuzzy Systems , 18–22doi:10.1109/fuzzy.1993.327470

  17. [26]

    AU-AIR: A Multi-modal Unmanned Aerial Vehicle Dataset for Low Altitude Traffic Surveillance URL: http://arxiv.org/abs/2001.11737

    Bozcan, I., Kayacan, E., 2020. AU-AIR: A Multi-modal Unmanned Aerial Vehicle Dataset for Low Altitude Traffic Surveillance URL: http://arxiv.org/abs/2001.11737

  18. [27]

    The EuroCity Persons Dataset: A Novel Benchmark for Object Detection

    Braun, M., Krebs, S., Flohr, F., Gavrila, D.M., 2018. The EuroCity Persons Dataset: A Novel Benchmark for Object Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 41, 1844–1861. URL: http://arxiv.org/ abs/1805.07193http://dx.doi.org/10.1109/TPAMI.2019.2...

  19. [28]

    Segmentation and Recognition using SfM Point Clouds

    Brostow, G., Shotton, J., Fauqueur, J., Cipolla, R., 2008a. Segmentation and Recognition using SfM Point Clouds. Eccv , 1–15

  20. [29]

    Segmentation and Recognition Using Structure from Motion Point Clouds , 44–57

    Brostow, G.J., Shotton, J., Fauqueur, J., Cipolla, R., 2008b. Segmentation and Recognition Using Structure from Motion Point Clouds , 44–57

  21. [30]

    Object Segmentation by Long Term Analysis of Point Trajectories, in: Daniilidis, K., Maragos, P., Paragios, N

    Brox, T., Malik, J., 2010. Object Segmentation by Long Term Analysis of Point Trajectories, in: Daniilidis, K., Maragos, P., Paragios, N. (Eds.), Computer Vision – ECCV 2010, Springer Berlin Heidelberg, Berlin, Heidelberg. pp. 282–295

  22. [31]

    The 2019 DA VIS Challenge on VOS: Unsupervised Multi-Object Segmentation , 1–4URL: http://arxiv.org/abs/1905.00737

    Caelles, S., Pont-Tuset, J., Perazzi, F., Montes, A., Maninis, K.K., Van Gool, L., 2019. The 2019 DA VIS Challenge on VOS: Unsupervised Multi-Object Segmentation , 1–4URL: http://arxiv.org/abs/1905.00737

  23. [32]

    Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O.,

  24. [33]

    Multi-Modality Vertebra Recognition in Arbitrary Views Using 3D Deformable Hierarchical Model

    Cai, Y., Osman, S., Sharma, M., Landis, M., Li, S., 2015. Multi-Modality Vertebra Recognition in Arbitrary Views Using 3D Deformable Hierarchical Model. IEEE Transactions on Medical Imaging 34, 1676–1693. doi: 10.1109/TMI. 2015.2392054

  25. [34]

    Caesar, H., Uijlings, J., Ferrari, V., 2018. COCO-Stuff Thing and Stuff Classes in Context - Caesar, Uijlings, Ferrari - 23 2016.pdf , 1209–1218URL: http://openaccess.thecvf.com/content{\_}cvpr{\_}2018/html/Caesar{\_}COCO-Stuff{\_ }Thing{\_}and{\_}CVPR{\_}2018{\_}paper.html

  26. [35]

    ISIC 2018

    Canfield, Kittler, H., Codella, N., Celebi, M.E., Dana, K., Halpern, A., Helba, B., Tschandl, P., . ISIC 2018. URL: https://challenge2018.isic-archive.com/

  27. [36]

    Nucleus segmentation across imaging experiments: the 2018 Data Science Bowl

    Caicedo, J.C., Goodman, A., Karhohs, K.W., Cimini, B.A., Ackerman, J., Haghighi, M., Heng, C., Becker, T., Doan, M., McQuin, C., Rohban, M., Singh, S., Carpenter, A.E., 2019. Nucleus segmentation across imaging experiments: the 2018 Data Science Bowl. Nature Methods 16, 1247–1...

  28. [37]

    Argoverse: 3D tracking and forecasting with rich maps

    Chang, M.F., Lambert, J., Sangkloy, P., Singh, J., Bak, S., Hartnett, A., Wang, D., Carr, P., Lucey, S., Ramanan, D., Hays, J., 2019. Argoverse: 3D tracking and forecasting with rich maps. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recog...

  29. [38]

    Cao, Q., Shen, L., Xie, W., Parkhi, O.M., Zisserman, A., 2018. VGGFace2: A dataset for recognising faces across pose and age, in: Proceedings - 13th IEEE International Conference on Automatic Face and Gesture Recognition, FG 2018, Institute of Electrical and Electronics Engine...

  30. [39]

    Dˆ2-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios URL: http://arxiv.org/abs/1904.01975

    Che, Z., Li, G., Li, T., Jiang, B., Shi, X., Zhang, X., Lu, Y., Wu, G., Liu, Y., Ye, J., 2019. Dˆ2-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios URL: http://arxiv.org/abs/1904.01975

  31. [40]

    Return of the devil in the details: Delving deep into convolutional nets

    Chatfield, K., Simonyan, K., Vedaldi, A., Zisserman, A., 2014. Return of the devil in the details: Delving deep into convolutional nets. BMVC 2014 - Proceedings of the British Machine Vision Conference 2014 , 1–11doi: 10.5244/c.28.6

  32. [41]

    Cross-Age Reference Coding for Age-Invariant Face Recognition and Retrieval , 768–783

    Chen, B.C., Chen, C.S., Hsu, W., 2014a. Cross-Age Reference Coding for Age-Invariant Face Recognition and Retrieval , 768–783

  33. [42]

    High Performance Convolutional Neural Networks for Document Processing, in: Lorette, G

    Chellapilla, K., Puri, S., Simard, P., 2006. High Performance Convolutional Neural Networks for Document Processing, in: Lorette, G. (Ed.), Tenth International Workshop on Frontiers in Handwriting Recognition, Suvisoft, La Baule (France)

  34. [43]

    Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs

    Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L., 2017b. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence 40, 834–848

  35. [44]

    Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs

    Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L., 2017a. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence 40, 834–848

  36. [45]

    TensorMask: A Foundation for Dense Object Segmentation

    Chen, X., Girshick, R., He, K., Doll´ ar, P., 2019. TensorMask: A Foundation for Dense Object Segmentation

  37. [46]

    Rethinking atrous convolution for semantic image segmenta- tion

    Chen, L.C., Papandreou, G., Schroff, F., Adam, H., 2017c. Rethinking atrous convolution for semantic image segmenta- tion. arXiv preprint arXiv:1706.05587

  38. [47]

    A survey on object detection in optical remote sensing images

    Cheng, G., Han, J., 2016. A survey on object detection in optical remote sensing images. doi: 10.1016/j.isprsjprs. 2016.03.014

  39. [48]

    Chen, X., Mottaghi, R., Liu, X., Fidler, S., Urtasun, R., Yuille, A., 2014b. Detect what you can: Detecting and representing objects using holistic models and body parts, in: Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, IEEE Computer Soci...

  40. [49]

    Salientshape: Group saliency in image collections

    Cheng, M.M., Mitra, N., Huang, X., Hu, S.M., 2013. Salientshape: Group saliency in image collections. The Visual Computer 30, 1–10. doi: 10.1007/s00371-013-0867-4

  41. [50]

    Remote Sensing Image Scene Classification: Benchmark and State of the Art

    Cheng, G., Han, J., Lu, X., 2017. Remote Sensing Image Scene Classification: Benchmark and State of the Art. Proceedings of the IEEE 105, 1865–1883. URL: http://arxiv.org/abs/1703.00121http://dx.doi.org/10.1109/JPROC. 2017.2675998, doi:10.1109/JPROC.2017.2675998

  42. [51]

    Xception: Deep learning with depthwise separable convolutions

    Chollet, F., 2017. Xception: Deep learning with depthwise separable convolutions. Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 2017-Janua, 1800–1807. doi: 10.1109/CVPR.2017.195

  43. [52]

    KAIST Multi-Spectral Day/Night Data Set for Autonomous and Assisted Driving

    Choi, Y., Kim, N., Hwang, S., Park, K., Yoon, J.S., An, K., Kweon, I.S., 2018. KAIST Multi-Spectral Day/Night Data Set for Autonomous and Assisted Driving. IEEE Transactions on Intelligent Transportation Systems 19, 934–948. doi:10.1109/TITS.2018.2791533

  44. [53]

    Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC) , 1–12

    Codella, N., Rotemberg, V., Tschandl, P., Celebi, M.E., Dusza, S., Gutman, D., Helba, B., Kalloo, A., Liopyris, K., Marchetti, M., Kittler, H., Halpern, A., 2019. Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collabor...

  45. [54]

    Functional Map of the World

    Christie, G., Fendley, N., Wilson, J., Mukherjee, R., 2017. Functional Map of the World. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , 6172–6180URL: http://arxiv.org/abs/ 1711.07846

  46. [55]

    The Cityscapes Dataset for Semantic Urban Scene Understanding

    Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., Schiele, B., 2016. The Cityscapes Dataset for Semantic Urban Scene Understanding. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 20...

  47. [56]

    A critical evaluation of the Next Generation Simulation (NGSIM) vehicle trajectory dataset

    Coifman, B., Li, L., 2017. A critical evaluation of the Next Generation Simulation (NGSIM) vehicle trajectory dataset. Transportation Research Part B: Methodological 105, 362–377. doi: 10.1016/j.trb.2017.09.018

  48. [57]

    SARAS-ESAD Dataset

    Cuzzolin, F., Bawa, V.S., Skarga-Bandurova, I., Singh, G., 2020b. SARAS-ESAD Dataset. URL: https://saras-esad. grand-challenge.org/Dataset/

  49. [58]

    SARAS-ESAD 2020

    Cuzzolin, F., Bawa, V.S., Skarga-Bandurova, I., Singh, G., 2020a. SARAS-ESAD 2020

  50. [59]

    EchoNet-Dynamic Dataset

    David, O., Bryan, H., Amirata, G., Matt P., L., Euan A., A., David H., L., James Y., Z., 2019. EchoNet-Dynamic Dataset

  51. [60]

    Histograms of oriented gradients for human detection

    Dalal, N., Triggs, B., 2005. Histograms of oriented gradients for human detection. Proceedings - 2005 IEEE Computer 24 Society Conference on Computer Vision and Pattern Recognition, CVPR 2005 I, 886–893. doi: 10.1109/CVPR.2005.177

  52. [61]

    ImageNet: A large-scale hierarchical image database , 248–255doi:10.1109/cvpr.2009.5206848

    Deng, J., Dong, W., Socher, R., Li, L.J., Kai Li, Li Fei-Fei, 2010. ImageNet: A large-scale hierarchical image database , 248–255doi:10.1109/cvpr.2009.5206848

  53. [62]

    DeepGlobe 2018: A Challenge to Parse the Earth through Satellite Images

    Demir, I., Koperski, K., Lindenbaum, D., Pang, G., Huang, J., Basu, S., Hughes, F., Tuia, D., Raskar, R., Works, C., . DeepGlobe 2018: A Challenge to Parse the Earth through Satellite Images. Technical Report

  54. [63]

    Automatic detection of geometrical anomalies in composites manufacturing : a deep learning-based computer vision approach

    Djavadifar, A., 2020. Automatic detection of geometrical anomalies in composites manufacturing : a deep learning-based computer vision approach. Ph.D. thesis

  55. [64]

    D2-City Detection Domain Adaptation Challenge

    DiDi, 2019. D2-City Detection Domain Adaptation Challenge

  56. [65]

    ELCAP Public Lung Image Database

    ELCAP, 2003. ELCAP Public Lung Image Database. URL: http://www.via.cornell.edu/lungdb.html

  57. [66]

    Pedestrian detection: A benchmark, Institute of Electrical and Electronics Engineers (IEEE)

    Dollar, P., Wojek, C., Schiele, B., Perona, P., 2010. Pedestrian detection: A benchmark, Institute of Electrical and Electronics Engineers (IEEE). pp. 304–311. doi: 10.1109/cvpr.2009.5206631

  58. [67]

    SpaceNet: A Remote Sensing Dataset and Challenge Series

    Etten, A.V., Lindenbaum, D., Bacastow, T., . SpaceNet: A Remote Sensing Dataset and Challenge Series. Technical Report

  59. [68]

    Monocular pedestrian detection: Survey and experiments, in: IEEE Transactions on Pattern Analysis and Machine Intelligence, pp

    Enzweiler, M., Gavrila, D.M., 2009. Monocular pedestrian detection: Survey and experiments, in: IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 2179–2195. doi: 10.1109/TPAMI.2008.260

  60. [69]

    ”Hello! My name is

    Everingham, M., Sivic, J., Zisserman, A., 2006. ”Hello! My name is... Buffy” - Automatic naming of characters in TV video. BMVC 2006 - Proceedings of the British Machine Vision Conference 2006 , 899–908

  61. [70]

    The Pascal Visual Object Classes Challenge: A Retrospective

    Everingham, M., Eslami, S.M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A., 2014. The Pascal Visual Object Classes Challenge: A Retrospective. International Journal of Computer Vision 111, 98–136. doi: 10.1007/ s11263-014-0733-5

  62. [71]

    Camouflaged Object Detection

    Fan, D.p., Guolei, G.p.J., Cheng, S.M.m., Shen, J., Shao, L., 2020a. Camouflaged Object Detection

  63. [72]

    The pascal visual object classes (VOC) challenge

    Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A., 2010. The pascal visual object classes (VOC) challenge. International Journal of Computer Vision 88, 303–338. doi: 10.1007/s11263-009-0275-4

  64. [73]

    Salient objects in clutter: Bringing salient object detection to the foreground

    Fan, D.P., Liu, J.J., Gao, S., Hou, Q., Borji, A., Cheng, M.M., 2018. Salient objects in clutter: Bringing salient object detection to the foreground. European Conference on Computer Vision (ECCV)

  65. [74]

    Camouflaged object detection , 2774–2784doi: 10

    Fan, D.P., Ji, G.P., Sun, G., Cheng, M.M., Shen, J., Shao, L., 2020b. Camouflaged object detection , 2774–2784doi: 10. 1109/CVPR42600.2020.00285

  66. [75]

    Learning Generative Visual Models from Few Training Examples :

    Fei- Fei, L., Fergus, R., Perona, P., 2004. Learning Generative Visual Models from Few Training Examples :. Conference on Computer Vision and Pattern Recognition Workshop (CVPR 2004) 00, 178. URL: http://dx.doi.org/10.1109/ CVPR.2004.109, doi:10.1109/CVPR.2004.109

  67. [76]

    JumpCut: Non-Successive Mask Transfer and Interpolation for Video Cutout

    Fan, Q., Zhong, F., Lischinski, D., Cohen-Or, D., Chen, B., 2015. JumpCut: Non-Successive Mask Transfer and Interpolation for Video Cutout. ACM Trans. Graph. 34. URL: https://doi.org/10.1145/2816795.2818105, doi: 10. 1145/2816795.2818105

  68. [78]

    WordNet: an Electronic Lexical Database

    Fellbaum, C., 1998. WordNet: an Electronic Lexical Database. Bradford Books

  69. [79]

    Construction of a Machine Learning Dataset through Collaboration: The RSNA 2019 Brain CT Hemorrhage Challenge

    Flanders, A.E., Prevedello, L.M., Shih, G., Halabi, S.S., Kalpathy-Cramer, J., Ball, R., Mongan, J.T., Stein, A., Kita- mura, f.C., Lungren, Mattew, P., Choudhary, G., Cala.lesley, Coelho, L., Mogensen, M., Moron, F., Miller, E., Ikuta, I., Zohrabian, V., Mcdonnell, O., Lincol...

  70. [80]

    Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges

    Feng, D., Haase-Sch¨ utz, C., Rosenbaum, L., Hertlein, H., Gl¨ aser, C., Timm, F., Wiesbeck, W., Dietmayer, K., 2021. Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges. IEEE Transactions on Intelligent Transpo...

  71. [81]

    A review on deep learning techniques applied to semantic segmentation

    Garcia-Garcia, A., Orts-Escolano, S., Oprea, S., Villena-Martinez, V., Garcia-Rodriguez, J., 2017. A review on deep learning techniques applied to semantic segmentation

  72. [82]

    Research and development of power grid dispatching operation control system based on transmission section control

    Gan, D., Lin, G., Wu, H., Peng, J., Zhang, Y., Liao, B., Huang, Y., Zheng, Q., Zhang, N., 2017. Research and development of power grid dispatching operation control system based on transmission section control. Dianli Xitong Baohu yu Kongzhi/Power System Protection and Control...

  73. [83]

    The KITTI 2D Object Evaluation Benchmark

    Geiger, A., Lenz, P., Stiller, C., Urtasun, R., a. The KITTI 2D Object Evaluation Benchmark

  74. [84]

    Ge, Y., Zhang, R., Wang, X., Tang, X., Luo, P., 2019. Deepfashion2: A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images, in: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, ...

  75. [85]

    Vision meets robotics: The KITTI dataset

    Geiger, A., Lenz, P., Stiller, C., Urtasun, R., 2013. Vision meets robotics: The KITTI dataset. International Journal of Robotics Research 32, 1231–1237. doi: 10.1177/0278364913491297

  76. [86]

    The KITTI 3D Object Evaluation Benchmark

    Geiger, A., Lenz, P., Stiller, C., Urtasun, R., b. The KITTI 3D Object Evaluation Benchmark

  77. [87]

    Fast r-cnn, in: Proceedings of the IEEE international conference on computer vision, pp

    Girshick, R., 2015. Fast r-cnn, in: Proceedings of the IEEE international conference on computer vision, pp. 1440–1448

  78. [89]

    Overview of LifeCLEF Plant identification task 2019: Diving into data deficient tropical countries, in: CEUR Workshop Proceedings, pp

    Go¨ eau, H., Bonnet, P., Joly, A., 2019. Overview of LifeCLEF Plant identification task 2019: Diving into data deficient tropical countries, in: CEUR Workshop Proceedings, pp. 9–12

  79. [90]

    Rich feature hierarchies for accurate object detection and semantic segmentation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

    Girshick, R., Donahue, J., Darrell, T., Malik, J., 2014. Rich feature hierarchies for accurate object detection and semantic segmentation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 580–587. 25

  80. [92]

    STARE Database

    Goldbaum, M., 1975. STARE Database

  81. [93]

    MS-Celeb-1M: A Dataset and Benchmark for Large-Scale Face Recognition

    Guo, Y., Zhang, L., Hu, Y., He, X., Gao, J., 2016. MS-Celeb-1M: A Dataset and Benchmark for Large-Scale Face Recognition

  82. [94]

    Caltech-256 Object Category Dataset , 300

    Griffin, Greg, 2007. Caltech-256 Object Category Dataset , 300

  83. [95]

    Semantic Contours from Inverse Detectors * - Hariharan et al.pdf

    Hariharan, B., Arbel´ aez, P., Bourdev, L., Maji, S., Malik, J., 2011. Semantic Contours from Inverse Detectors * - Hariharan et al.pdf. International Conference on Computer Vision , 8URL: http://home.bharathh.info/pubs/pdfs/ BharathICCV2011.pdf

  84. [96]

    Lvis: A dataset for large vocabulary instance segmentation

    Gupta, A., Dollar, P., Girshick, R., 2019. Lvis: A dataset for large vocabulary instance segmentation. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2019-June, 5351–5359. doi: 10.1109/ CVPR.2019.00550

  85. [97]

    Deep residual learning for image recognition, in: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE Computer Society

    He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep residual learning for image recognition, in: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE Computer Society. pp. 770–778. URL: http://image-net.org/challenges/LSVRC/2015/, do...

  86. [98]

    Mask r-cnn, in: Proceedings of the IEEE international conference on computer vision, pp

    He, K., Gkioxari, G., Doll´ ar, P., Girshick, R., 2017. Mask r-cnn, in: Proceedings of the IEEE international conference on computer vision, pp. 2961–2969

  87. [99]

    Philip Chang, K., Munishkumaran, S., 2001

    Heath, M., Bowyer, K., Kopans, D., Morre, R., Kegelmeyer, W. Philip Chang, K., Munishkumaran, S., 2001. THE DIGITAL DATABASE FOR SCREENING MAMMOGRAPHY. Medical Physics Publishing

  88. [100]

    Philip Chang, K., Munishkumaran, S.,

    Heath, M., Bowyer, K., Kopans, D., Morre, R., Kegelmeyer, W. Philip Chang, K., Munishkumaran, S., . Current Status of the Digital Database for Screening Mammography. Digital Mammography , 457–460doi: https://doi.org/10.1007/ 978-94-011-5318-8_75

  89. [101]

    The KiTS19 Challenge Data: 300 Kidney Tumor Cases with Clinical Context, CT Semantic Segmentations, and Surgical Outcomes , 1–14

    Heller, N., Sathianathen, N., Kalapara, A., Walczak, E., Moore, K., Kaluzniak, H., Rosenberg, J., Blake, P., Rengel, Z., Oestreich, M., Dean, J., Tradewell, M., Shah, A., Tejpaul, R., Edgerton, Z., Peterson, M., Raza, S., Regmi, S., Papanikolopoulos, N., Weight, C., 2019. The ...

  90. [102]

    Learning Spatial Context: Using Stuff to Find Things

    Heitz, G., Koller, D., . Learning Spatial Context: Using Stuff to Find Things. Technical Report

  91. [103]

    Reducing the dimensionality of data with neural networks

    Hinton, G.E., Salakhutdinov, R.R., 2006. Reducing the dimensionality of data with neural networks. science 313, 504–507

  92. [104]

    A fast learning algorithm for deep belief nets

    Hinton, G.E., Osindero, S., Teh, Y.W., 2006. A fast learning algorithm for deep belief nets. Neural computation 18, 1527–1554

  93. [105]

    The iNaturalist Species Classification and Detection Dataset, in: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp

    Horn, G.V., Aodha, O.M., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., Belongie, S., 2018. The iNaturalist Species Classification and Detection Dataset, in: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 876...

  94. [106]

    A DATA-DRIVEN APPROACH TO CLEANING LARGE F ACE DATASETS

    Hong-Wei, N., Stefan, W., 2014. A DATA-DRIVEN APPROACH TO CLEANING LARGE F ACE DATASETS. Inter- national Conference on Image Processing(ICIP) , 343–347

  95. [107]

    Hosseini, M.S., Chan, L., Tse, G., Tang, M., Deng, J., Norouzi, S., Rowsell, C., Plataniotis, K.N., Damaskinos, S., 2019. Atlas of digital pathology: A generalized hierarchical histological tissue type-annotated database for deep learning, in: Proceedings of the IEEE Computer ...

  96. [108]

    Building a bird recognition app and large scale dataset with citizen scientists : The fine print in fine-grained dataset collection

    Horn, G.V., Branson, S., Farrell, R., Barry, J., Tech, C., . Building a bird recognition app and large scale dataset with citizen scientists : The fine print in fine-grained dataset collection

  97. [109]

    Cross-domain image retrieval with a dual attribute-aware ranking network, in: Proceedings of the IEEE International Conference on Computer Vision, pp

    Huang, J., Feris, R., Chen, Q., Yan, S., 2015. Cross-domain image retrieval with a dual attribute-aware ranking network, in: Proceedings of the IEEE International Conference on Computer Vision, pp. 1062–1070. doi: 10.1109/ICCV.2015.127

  98. [110]

    Huang, G.B., Mattar, M., Berg, T., Learned-Miller, E., Learned-Miller, E., 2008. Labeled Faces in the Wild: A Database forStudying Face Recognition in Unconstrained Environments Labeled Faces in the Wild: A Database for Studying Face Recognition in Unconstrained Environments. ...

  99. [111]

    LUNA 2016

    Jacobs, C., Setio, A.A.A., Traverso, A., Ginneken, B.V., 2016. LUNA 2016

  100. [112]

    CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison

    Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., Seekins, J., Mong, D.A., Halabi, S.S., Sandberg, J.K., Jones, R., Larson, D.B., Langlotz, C.P., Patel, B.N., Lungren, M.P., Ng, A.Y., 2019. CheXpert: A L...

  101. [113]

    Robust face detection using the Hausdorff distance

    Jesorsky, O., Kirchberg, K.J., Frischholz, R.W., 2001. Robust face detection using the Hausdorff distance. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 2091, 90–95. doi: 10.1007/3-540-45344-x_14

  102. [114]

    Supervoxel-Consistent Foreground Propagation in Video, pp

    Jain, S., Grauman, K., 2014. Supervoxel-Consistent Foreground Propagation in Video, pp. 656–671. doi: 10.1007/ 978-3-319-10593-2_43

  103. [115]

    CVPR 2018 W AD Video Segmentation Challenge

    Kaggle, 2018. CVPR 2018 W AD Video Segmentation Challenge. doi: https://www.kaggle.com/c/ cvpr-2018-autonomous-driving

  104. [116]

    The FERET evaluation methodology for face-recognition algorithms

    Jonathon Phillips, P., Moon, H., Rizvi, S.A., Rauss, P.J., 2000. The FERET evaluation methodology for face-recognition algorithms. IEEE Transactions on Pattern Analysis and Machine Intelligence 22, 1090–1104. doi: 10.1109/34.879790

  105. [117]

    FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age

    K¨ arkk¨ ainen, K., Joo UCLA, J., . FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age. Technical 26 Report. URL: https://github.com/joojs/fairface{\%}7D

  106. [118]

    Dstl satelite imagery feature detection

    Kaggle.com, 2017. Dstl satelite imagery feature detection. URL: https://www.kaggle.com/c/ dstl-satellite-imagery-feature-detection

  107. [119]

    The MegaFace benchmark: 1 million faces for recognition at scale

    Kemelmacher-Shlizerman, I., Seitz, S.M., Miller, D., Brossard, E., 2016. The MegaFace benchmark: 1 million faces for recognition at scale. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2016-Decem, 4873–4882. doi: 10.1109/CVPR.2016.527

  108. [120]

    The DIARETDB1 diabetic retinopathy database and evaluation protocol

    Kauppi, T., Kalesnykiene, V., Kamarainen, J.K., Lensu, L., Sorri, I., Raninen, A., Voutilainen, R., Pietil¨ a, J., K¨ alvi¨ ainen, H., Uusitalo, H., 2007. The DIARETDB1 diabetic retinopathy database and evaluation protocol. BMVC 2007 - Pro- ceedings of the British Machine Visi...

  109. [121]

    AnimalWeb: A Large-Scale Hierarchical Dataset of Annotated Animal Faces , 1–15URL: http://arxiv.org/abs/1909

    Khan, M.H., McDonagh, J., Khan, S., Shahabuddin, M., Arora, A., Khan, F.S., Shao, L., Tzimiropoulos, G., 2019. AnimalWeb: A Large-Scale Hierarchical Dataset of Annotated Animal Faces , 1–15URL: http://arxiv.org/abs/1909. 04951

  110. [122]

    Lyft Level 5 A V Dataset

    Kesten, R., Usman, M., Houston, J., Pandya, T., Nadhamuni, K., Ferreira, A., Yuan, M., Low, B., Jain, A., Ondruska, P., Omari, S., Shah, S., Kulkarni, A., Kazakova, A., Tao, C., Platinsky, L., Jiang, W., Shet., V., 2019. Lyft Level 5 A V Dataset. URL: https://level5.lyft.com/dataset/

  111. [123]

    Where to buy it: Matching street clothing photos in online shops, in: Proceedings of the IEEE International Conference on Computer Vision, pp

    Kiapour, M.H., Han, X., Lazebnik, S., Berg, A.C., Berg, T.L., 2015. Where to buy it: Matching street clothing photos in online shops, in: Proceedings of the IEEE International Conference on Computer Vision, pp. 3343–3351. doi:10.1109/ICCV.2015.382

  112. [124]

    Novel dataset for fine-grained image categorization

    Khosla, A., Jayadevaprakash, N., Yao, B., Fei-Fei, L., 2011. Novel dataset for fine-grained image categorization. Proc. IEEE Conf. Comput. Vision and Pattern Recognition

  113. [125]

    The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for Validation of Highly Automated Driving Systems

    Krajewski, R., Bock, J., Kloeker, L., Eckstein, L., 2018. The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for Validation of Highly Automated Driving Systems. IEEE Conference on Intelligent Transportation Systems, Proceedings, ITSC 201...

  114. [126]

    Klare, B.F., Klein, B., Taborsky, E., Blanton, A., Cheney, J., Allen, K., Grother, P., Mah, A., Burge, M., Jain, A.K.,

  115. [127]

    Learning Multiple Layers of Features from Tiny Images

    Krizhevsky, A., 2012. Learning Multiple Layers of Features from Tiny Images. University of Toronto

  116. [128]

    Imagenet classification with deep convolutional neural networks

    Krizhevsky, A., Sutskever, I., Hinton., G.E., 2012. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems , 1097–1105URL: http://arxiv.org/abs/1102.0183

  117. [129]

    Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

    Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., Bernstein, M.S., Fei-Fei, L., 2017. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. International Journal of Compute...

  118. [130]

    The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale , 1–20URL: http://arxiv.org/abs/1811.00982

    Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., Duerig, T., Ferrari, V., 2018. The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at s...

  119. [131]

    xView: Objects in Context in Overhead Imagery URL: http://arxiv.org/abs/1802.07856

    Lam, D., Kuzma, R., McGee, K., Dooley, S., Laielli, M., Klaric, M., Bulatov, Y., McCord, B., 2018. xView: Objects in Context in Overhead Imagery URL: http://arxiv.org/abs/1802.07856

  120. [132]

    Attribute and simile classifiers for face verification

    Kumar, N., Berg, A.C., Belhumeur, P.N., Nayar, S.K., 2009. Attribute and simile classifiers for face verification. Pro- ceedings of the IEEE International Conference on Computer Vision , 365–372doi: 10.1109/ICCV.2009.5459250

  121. [133]

    OASIS-3: Longitudinal Neuroimaging, Clinical, and Cognitive Dataset for Normal Aging and Alzheimer Disease

    LaMontagne, P.J., Benzinger, T.L., Morris, J.C., Keefe, S., Hornbeck, R., Xiong, C., Grant, E., Hassenstab, J., Moulder, K., Vlassenko, A., Raichle, Marcus, E., Carlos, C., Marcus, D., 2019. OASIS-3: Longitudinal Neuroimaging, Clinical, and Cognitive Dataset for Normal Aging a...

  122. [134]

    Survey on semantic segmentation using deep learning techniques

    Lateef, F., Ruichek, Y., 2019. Survey on semantic segmentation using deep learning techniques. Neurocomputing 338, 321–348. URL: https://www.sciencedirect.com/science/article/pii/S092523121930181X, doi: https://doi.org/10. 1016/j.neucom.2019.02.003

  123. [135]

    SegTHOR: Segmentation of Thoracic Organs at Risk in CT images , 1–16

    Lambert, Z., Petitjean, C., Dubray, B., Ruan, S., 2019. SegTHOR: Segmentation of Thoracic Organs at Risk in CT images , 1–16

  124. [136]

    Anabranch network for camouflaged object segmentation

    Le, T.N., Nguyen, T., Nie, Z., Tran, M.T., Sugimoto, A., 2019. Anabranch network for camouflaged object segmentation. Computer Vision and Image Understanding 184. doi: 10.1016/j.cviu.2019.04.006

  125. [137]

    Backpropagation applied to handwritten zip code recognition

    Lecun, Y., Boser, B., Denker, J.S., Henderson, D., Howard, R.E., Hubbard, W., Jackel, L.D., 1989. Backpropagation applied to handwritten zip code recognition. Neural computation 1, 541–551

  126. [138]

    Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories

    Lazebnik, S., Schmid, C., Ponce, J., 2006. Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2, 2169–2178. doi: 10.1109/CVPR.2006.68

  127. [139]

    Lecun, Y., Bottou, L., Bengio, Y., Ha, P., 1998. LeNet. Proceedings of the IEEE , 1–46doi: 10.1109/5.726791

  128. [140]

    Handwritten Digit Recognition with a Back-Propagation Network

    Lecun, Y., Others, 1997. Handwritten Digit Recognition with a Back-Propagation Network. Neural Information Pro- cessing Systems 2

  129. [141]

    Handwritten digit recognition with a back-propagation network, in: Advances in neural information processing systems, pp

    Lecun, Y., Boser, B.E., Denker, J.S., Henderson, D., Howard, R.E., Hubbard, W.E., Jackel, L.D., 1990. Handwritten digit recognition with a back-propagation network, in: Advances in neural information processing systems, pp. 396–404

  130. [142]

    Video Segmentation by Tracking Many Figure-Ground Segments, in: 2013 IEEE International Conference on Computer Vision, pp

    Li, F., Kim, T., Humayun, A., Tsai, D., Rehg, J.M., 2013. Video Segmentation by Tracking Many Figure-Ground Segments, in: 2013 IEEE International Conference on Computer Vision, pp. 2192–2199. doi: 10.1109/ICCV.2013.273

  131. [143]

    Visual saliency based on multiscale deep features doi: 10.1109/CVPR.2015.7299184

    Li, G., Yu, Y., 2015. Visual saliency based on multiscale deep features doi: 10.1109/CVPR.2015.7299184

  132. [144]

    LERA- Lower Extremity RAdiographs

    LERA, 2018. LERA- Lower Extremity RAdiographs. URL: https://aimi.stanford.edu/ lera-lower-extremity-radiographs-2

  133. [145]

    StructSeg 2019

    Li, H., Zhou, J., Deng, J., Chen, M., SenseTime, YINO, Zhejiang Cancer Hospital, 2019. StructSeg 2019

  134. [146]

    A review of remote sensing image classification techniques: the role of spatio-contextual information

    Li, M., Zang, S., Zhang, B., Li, S., Wu, C., 2014a. A review of remote sensing image classification techniques: the role of spatio-contextual information. European Journal of Remote Sensing 47, 389–411. URL: https://doi.org/10.5721/ EuJRS20144723, doi:10.5721/EuJRS20144723

  135. [147]

    Automatic Structure Segmentation for Radiotherapy Planning Challenge 2020

    Li, H., Chen, M., 2020. Automatic Structure Segmentation for Radiotherapy Planning Challenge 2020. doi: 10.5281/ 27 zenodo.3718885

  136. [148]

    Multi-scale cascade network for salient object detection , 439–447doi:10.1145/3123266.3123290

    Li, X., Yang, F., Cheng, H., Chen, J., Guo, Y., Chen, L., 2017. Multi-scale cascade network for salient object detection , 439–447doi:10.1145/3123266.3123290

  137. [149]

    Li, X., Yang, F., Cheng, H., Liu, W., Shen, D., 2018. Contour knowledge transfer for salient object detec- tion: 15th european conference, munich, germany, september 8-14, 2018, proceedings, part xv , 370–385doi: 10.1007/ 978-3-030-01267-0_22

  138. [150]

    Li, S., Wang, 2019. AASCE. URL: https://aasce19.grand-challenge.org/

  139. [151]

    Feature pyramid networks for object detection

    Lin, T.Y., Doll´ ar, P., Girshick, R., He, K., Hariharan, B., Belongie, S., 2016. Feature pyramid networks for object detection

  140. [152]

    Microsoft COCO: Common objects in context

    Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll´ ar, P., Zitnick, C.L., 2014. Microsoft COCO: Common objects in context. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformat...

  141. [153]

    The secrets of salient object segmentation

    Li, Y., Hou, X., Koch, C., Rehg, J., Yuille, A., 2014b. The secrets of salient object segmentation. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition doi: 10.1109/CVPR.2014.43

  142. [154]

    Fast Multiclass Vehicle Detection on Aerial Images

    Liu, K., Mattyus, G., 2015. Fast Multiclass Vehicle Detection on Aerial Images. IEEE Geoscience and Remote Sensing Letters 12, 1938–1942. doi: 10.1109/LGRS.2015.2439517

  143. [156]

    Nonparametric scene parsing via label transfer

    Liu, C., Yuen, J., Torralba, A., 2015a. Nonparametric scene parsing via label transfer. Dense Image Correspondences for Computer Vision 33, 207–236. doi: 10.1007/978-3-319-23048-1_10

  144. [157]

    Learning to detect a salient object , 1–8doi: 10.1109/CVPR

    Liu, T., Sun, J., Zheng, N.N., Tang, X., Shum, H.Y., 2007. Learning to detect a salient object , 1–8doi: 10.1109/CVPR. 2007.383047

  145. [158]

    SSD: Single Shot MultiBox Detector doi:10.1007/978-3-319-46448-0_2

    Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., Berg, A.C., 2015b. SSD: Single Shot MultiBox Detector doi:10.1007/978-3-319-46448-0_2

  146. [159]

    Deep learning for generic object detection: A survey

    Liu, L., Ouyang, W., Wang, X., Fieguth, P., Chen, J., Liu, X., Pietikinen, M., 2020b. Deep learning for generic object detection: A survey. International Journal of Computer Vision 128, 261–318. URL: https://doi.org/10.1007/ s11263-019-01247-4 , doi:10.1007/s11263-019-01247-4

  147. [160]

    Deep learning face attributes in the wild

    Liu, Z., Luo, P., Wang, X., Tang, X., 2015c. Deep learning face attributes in the wild. Proceedings of the IEEE International Conference on Computer Vision 2015 Inter, 3730–3738. doi: 10.1109/ICCV.2015.425

  148. [161]

    Object recognition from local scale-invariant features, in: Proceedings of the IEEE International Conference on Computer Vision, IEEE

    Lowe, D.G., 1999. Object recognition from local scale-invariant features, in: Proceedings of the IEEE International Conference on Computer Vision, IEEE. pp. 1150–1157. doi: 10.1109/iccv.1999.790410

  149. [162]

    Liu, Z., Luo, P., Qiu, S., Wang, X., Tang, X., 2016. DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations, in: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 1096–1104. doi: 10.1109/CVPR.2016.124

  150. [163]

    1 Year , 1000km : The Oxford RobotCar Dataset 3

    Maddern, W., Pascoe, G., Linegar, C., Newman, P., . 1 Year , 1000km : The Oxford RobotCar Dataset 3

  151. [164]

    SMIR Database URL: https://www.smir.ch

    Maier, O., 2015. SMIR Database URL: https://www.smir.ch

  152. [165]

    Lyft 3D Object Detection for Autonomous Vehicles

    Lyft, 2019. Lyft 3D Object Detection for Autonomous Vehicles. URL: https://www.kaggle.com/c/ 3d-object-detection-for-autonomous-vehicles

  153. [166]

    The AR face database

    Martinez, A.M., 1998. The AR face database. CVC Technical Report24

  154. [167]

    Deep Face Recognition: A Survey

    Masi, I., Wu, Y., Hassner, T., Natarajan, P., 2019. Deep Face Recognition: A Survey. Proceedings - 31st Conference on Graphics, Patterns and Images, SIBGRAPI 2018 , 471–478doi: 10.1109/SIBGRAPI.2018.00067

  155. [168]

    A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics

    Martin, D., Fowlkes, C., Tal, D., Malik, J., 2001. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. Proceedings of the IEEE International Conference on Computer Vision 2, 416–423. doi: 1...

  156. [169]

    Diversity in Faces , 1–29URL: http://arxiv.org/abs/1901.10436

    Merler, M., Ratha, N., Feris, R.S., Smith, J.R., 2019. Diversity in Faces , 1–29URL: http://arxiv.org/abs/1901.10436

  157. [170]

    Automotive radar dataset for deep learning based 3D object detection

    Meyer, M., Kuschk, G., 2019. Automotive radar dataset for deep learning based 3D object detection. EuRAD 2019 - 2019 16th European Radar Conference , 129–132

  158. [171]

    IARPA janus benchmark-C: Face dataset and protocol

    Maze, B., Adams, J., Duncan, J.A., Kalka, N., Miller, T., Otto, C., Jain, A.K., Niggel, W.T., Anderson, J., Cheney, J., Grother, P., 2018. IARPA janus benchmark-C: Face dataset and protocol. Proceedings - 2018 International Conference on Biometrics, ICB 2018 , 158–165doi: 10.1...

  159. [172]

    A Large Contextual Dataset for Classification, Detection and Counting of Cars with Deep Learning

    Mundhenk, T.N., Konjevod, G., Sakla, W.A., Boakye, K., 2016. A Large Contextual Dataset for Classification, Detection and Counting of Cars with Deep Learning. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in B...

  160. [173]

    National Library of Medicine, 2006. MedPix. URL: https://medpix.nlm.nih.gov/home

  161. [174]

    The role of context for object detection and semantic segmentation in the wild

    Mottaghi, R., Chen, X., Liu, X., Cho, N.G., Lee, S.W., Fidler, S., Urtasun, R., Yuille, A., 2014. The role of context for object detection and semantic segmentation in the wild. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , 89...

  162. [175]

    Columbia Object Image Library (COIL-100)

    Nene, S., Nayar, S., Murase, H., 1996a. Columbia Object Image Library (COIL-100). Technical Report 95, 223–303. URL: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.54.5914

  163. [176]

    Columbia Object Image Library (COIL-20)

    Nene, S., Nayar, S., Murase, H., 1996b. Columbia Object Image Library (COIL-20). Technical Report 95, 223–303. URL: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.54.5914

  164. [177]

    Level Playing Field for Million Scale Face Recognition

    Nech, A., Kemelmacher-Shlizerman, I., Allen, P.G., . Level Playing Field for Million Scale Face Recognition. Technical Report. 28

  165. [178]

    Neumann, L., Karg, M., Zhang, S., Scharfenberger, C., Piegert, E., Mistr, S., Prokofyeva, O., Thiel, R., Vedaldi, A., Zisserman, A., Schiele, B., 2019. NightOwls: A Pedestrians at Night Dataset, in: Lecture Notes in Computer Science (including subseries Lecture Notes in Artifi...

  166. [179]

    Automated flower classification over a large number of classes

    Nilsback, M.E., Zisserman, A., 2008. Automated flower classification over a large number of classes. Proceedings - 6th Indian Conference on Computer Vision, Graphics and Image Processing, ICVGIP 2008 , 722–729doi: 10.1109/ICVGIP. 2008.47

  167. [180]

    The Mapillary Vistas Dataset for Semantic Understanding of Street Scenes

    Neuhold, G., Ollmann, T., Bulo, S.R., Kontschieder, P., 2017. The Mapillary Vistas Dataset for Semantic Understanding of Street Scenes. Proceedings of the IEEE International Conference on Computer Vision 2017-Octob, 5000–5009. doi: 10. 1109/ICCV.2017.534

  168. [181]

    Odir, 2019. ODIR-5K. URL: http://www.kaggle.com/andrewmvd/ocular-disease-recognition-odir5k

  169. [182]

    Orlando, J.I., Fu, H., Breda, J.B., van Keer, K., Bathula, D.R., Diaz-Pinto, A., Fang, R., Heng, P., Kim, J., Lee, J., Lee, J., Li, X., Liu, P., Lu, S., Murugesan, B., Naranjo, V., Phaye, S.S.R., Shankaranarayana, S.M., Sikka, A., Son, J., van den Hengel, A., Wang, S., Wu, J.,...

  170. [183]

    Segmentation of Moving Objects by Long Term Video Analysis

    Ochs, P., Malik, J., Brox, T., 2014. Segmentation of Moving Objects by Long Term Video Analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence 36, 1187–1200. doi: 10.1109/TPAMI.2013.242

  171. [184]

    Trainable system for object detection

    Papageorgiou, C., Poggio, T., 2000. Trainable system for object detection. International Journal of Computer Vision 38, 15–33. doi: 10.1023/A:1008162616689

  172. [185]

    Deep Face Recognition , 41.1–41.12doi: 10.5244/c.29.41

    Parkhi, O.M., Vedaldi, A., Zisserman, A., 2015. Deep Face Recognition , 41.1–41.12doi: 10.5244/c.29.41

  173. [186]

    CoRR abs/1910.03667

    REFUGE challenge: A unified framework for evaluating automated methods for glaucoma assessment from fundus photographs. CoRR abs/1910.03667

  174. [187]

    Training support vector machines: An application to face detection

    Osuna, E., Freund, R., Girosi, F., 1997. Training support vector machines: An application to face detection. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , 130–136doi: 10.1109/cvpr. 1997.609310

  175. [188]

    A*3D Dataset: Towards Autonomous Driving in Challenging Environments

    Pham, Q.H., Sevestre, P., Pahwa, R.S., Zhan, H., Pang, C.H., Chen, Y., Mustafa, A., Chandrasekhar, V., Lin, J., 2019. A*3D Dataset: Towards Autonomous Driving in Challenging Environments

  176. [189]

    The FERET database and evaluation procedure for face- recognition algorithms

    Phillips, P.J., Wechsler, H., Huang, J., Rauss, P.J., 1998. The FERET database and evaluation procedure for face- recognition algorithms. Image and Vision Computing 16, 295–306. doi: 10.1016/s0262-8856(97)00070-x

  177. [190]

    The H3D dataset for full-surround 3D multi-object detection and tracking in crowded urban scenes

    Patil, A., Malla, S., Gang, H., Chen, Y.T., 2019. The H3D dataset for full-surround 3D multi-object detection and tracking in crowded urban scenes. Proceedings - IEEE International Conference on Robotics and Automation 2019-May, 9552–9557. doi: 10.1109/ICRA.2019.8793925

  178. [191]

    SUN attribute database: Discovering, annotating, and recognizing scene attributes

    Patterson, G., Hays, J., 2012. SUN attribute database: Discovering, annotating, and recognizing scene attributes. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , 2751–2758doi: 10. 1109/CVPR.2012.6247998

  179. [192]

    RSNA Intracranial Hemorrhage Detection

    Radiological Society of North America, 2019. RSNA Intracranial Hemorrhage Detection

  180. [193]

    MURA: Large Dataset for Abnormality Detection in Musculoskeletal Radiographs , 1–10

    Rajpurkar, P., Irvin, J., Bagul, A., Ding, D., Duan, T., Mehta, H., Yang, B., Zhu, K., Laird, D., Ball, R.L., Langlotz, C., Shpanskaya, K., Lungren, M.P., Ng, A.Y., 2017. MURA: Large Dataset for Abnormality Detection in Musculoskeletal Radiographs , 1–10

  181. [194]

    Learning object class detectors from weakly annotated video

    Prest, A., Leistner, C., Civera, J., Schmid, C., Ferrari, V., 2012. Learning object class detectors from weakly annotated video. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , 3282– 3289doi:10.1109/CVPR.2012.6248065

  182. [195]

    Recognizing indoor scenes

    Quattoni, A., Torralba, A., 2010. Recognizing indoor scenes. 2009 IEEE Conference on Computer Vision and Pattern Recognition , 413–420doi: 10.1109/cvpr.2009.5206537

  183. [196]

    Vehicle detection in aerial imagery: A small target detection benchmark

    Razakarivony, S., Jurie, F., 2016. Vehicle detection in aerial imagery: A small target detection benchmark. Journal of Visual Communication and Image Representation 34, 187–203. doi: 10.1016/j.jvcir.2015.11.002

  184. [197]

    Real, E., Shlens, J., Mazzocchi, S., Pan, X., Vanhoucke, V., 2017. YouTube-BoundingBoxes: A large high-precision human-annotated data set for object detection in video, in: Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, pp. 7464–7473....

  185. [198]

    Unsupervised Learning of Invariant Feature Hierarchies with Applications to Object Recognition, in: 2007 IEEE Conference on Computer Vision and Pattern Recognition, pp

    Ranzato, M., Huang, F.J., Boureau, Y., LeCun, Y., 2007. Unsupervised Learning of Invariant Feature Hierarchies with Applications to Object Recognition, in: 2007 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1–8. doi:10.1109/CVPR.2007.383157

  186. [199]

    Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review

    Rawat, W., Wang, Z., 2017. Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review. Neural Computation 29, 2352–2449. doi: 10.1162/neco_a_00990

  187. [200]

    YOLOv3: An Incremental Improvement

    Redmon, J., Farhadi, A., 2018. YOLOv3: An Incremental Improvement

  188. [201]

    CuRIOUS 2019

    Reinertsen, I., Xiao, Y., Rivaz, H., Chabanas, M., 2019. CuRIOUS 2019. URL: https://curious2019.grand-challenge. org/

  189. [202]

    You Only Look Once: Unified, Real-Time Object Detection

    Redmon, J., Divvala, S., Girshick, R., Farhadi, A., . You Only Look Once: Unified, Real-Time Object Detection

  190. [203]

    YOLO9000: Better, Faster, Stronger

    Redmon, J., Farhadi, A., 2016. YOLO9000: Better, Faster, Stronger

  191. [204]

    Deep expectation of real and apparent age from a single image without facial landmarks Real age 20 years DEX age predic3on

    Rothe, R., Timofte, R., Van Gool, L., . Deep expectation of real and apparent age from a single image without facial landmarks Real age 20 years DEX age predic3on. Technical Report

  192. [205]

    Deep Expectation of Real and Apparent Age from a Single Image Without Facial Landmarks

    Rothe, R., Timofte, R., Van Gool, L., 2018. Deep Expectation of Real and Apparent Age from a Single Image Without Facial Landmarks. International Journal of Computer Vision 126, 144–157. doi: 10.1007/s11263-016-0940-3

  193. [206]

    Faster r-cnn: Towards real-time object detection with region proposal networks, in: Advances in neural information processing systems, pp

    Ren, S., He, K., Girshick, R., Sun, J., 2015. Faster r-cnn: Towards real-time object detection with region proposal networks, in: Advances in neural information processing systems, pp. 91–99

  194. [207]

    U-net: Convolutional networks for biomedical image segmentation, in: 29 International Conference on Medical image computing and computer-assisted intervention, pp

    Ronneberger, O., Fischer, P., Brox, T., 2015. U-net: Convolutional networks for biomedical image segmentation, in: 29 International Conference on Medical image computing and computer-assisted intervention, pp. 234–241

  195. [208]

    FaceNet: A Unified Embedding for Face Recognition and Clustering

    Schroff, F., Philbin, J., . FaceNet: A Unified Embedding for Face Recognition and Clustering. Technical Report

  196. [209]

    SEMANTIC3D

    Sensing, R., Sciences, S.I., Hackel, T., Savinov, N., Ladicky, L., Wegner, J.D., Schindler, K., Pollefeys, M., 2017. SEMANTIC3D . NET : A NEW LARGE-SCALE POINT CLOUD CLASSIFICATION IV, 6–9. doi: 10.5194/ isprs-annals-IV-1-W1-91-2017

  197. [210]

    ImageNet Large Scale Visual Recognition Challenge

    Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L., 2015. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision 115, 211–252. doi: 10.1007/s11263...

  198. [211]

    LabelMe: A database and web-based tool for image annotation

    Russell, B.C., Torralba, A., Murphy, K.P., Freeman, W.T., 2008. LabelMe: A database and web-based tool for image annotation. International Journal of Computer Vision 77, 157–173. doi: 10.1007/s11263-007-0090-8

  199. [212]

    TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation, in: Leonardis, A., Bischof, H., Pinz, A

    Shotton, J., Winn, J., Rother, C., Criminisi, A., 2006. TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation, in: Leonardis, A., Bischof, H., Pinz, A. (Eds.), Computer Vision – ECCV 2006, Springer Berlin Heidelberg, Berl...

  200. [213]

    Indoor segmentation and support inference from RGBD images

    Silberman, N., Hoiem, D., Kohli, P., Fergus, R., 2012. Indoor segmentation and support inference from RGBD images. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 7576 LNCS, 746–760. doi: 10.1...

  201. [214]

    Fast YOLO: A Fast You Only Look Once System for Real-time Embedded Object Detection in Video

    Shafiee, M.J., Chywl, B., Li, F., Wong, A., 2017. Fast YOLO: A Fast You Only Look Once System for Real-time Embedded Object Detection in Video

  202. [215]

    Objects365: A Large- scale, High-quality Dataset for Object Detection

    Shao, S., Li, Z., Zhang, T., Peng, C., Yu, G., Zhang, X., Li, J., Sun, J., Technology, M., 2019. Objects365: A Large- scale, High-quality Dataset for Object Detection. Proc. IEEE International Conference on Computer Vision (ICCV) , 8430–8439

  203. [216]

    SUN RGB-D: A RGB-D scene understanding benchmark suite

    Song, S., Lichtenberg, S.P., Xiao, J., 2015. SUN RGB-D: A RGB-D scene understanding benchmark suite. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 07-12-June, 567–576. doi: 10. 1109/CVPR.2015.7298655

  204. [217]

    Quantitative analysis of pulmonary emphysema using local binary patterns

    Sørensen, L., Shaker, S.B., De Bruijne, M., 2010. Quantitative analysis of pulmonary emphysema using local binary patterns. IEEE Transactions on Medical Imaging 29, 559–569. doi: 10.1109/TMI.2009.2038575

  205. [218]

    The CMU Pose, Illumination, and Expression (PIE) database

    Sim, T., Baker, S., Bsat, M., 2002. The CMU Pose, Illumination, and Expression (PIE) database. Proceedings - 5th IEEE International Conference on Automatic Face Gesture Recognition, FGR 2002 , 53–58doi: 10.1109/AFGR.2002.1004130

  206. [220]

    Sun, M., Yuan, Y., Zhou, F., Ding, E., 2018. Multi-Attention Multi-Class Constraint for Fine-grained Image Recognition, in: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), pp. 834–850. doi: 1...

  207. [221]

    Scalability in Perception for Autonomous Driving: Waymo Open Dataset

    Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., Vasudevan, V., Han, W., Ngiam, J., Zhao, H., Timofeev, A., Ettinger, S., Krivokon, M., Gao, A., Joshi, A., Zhang, Y., Shlens, J., Chen, Z., Anguelov, D., 2019. ...

  208. [222]

    An open, multi-vendor, multi-field-strength brain MR dataset and analysis of publicly available skull stripping methods agreement

    Souza, R., Lucena, O., Garrafa, J., Gobbi, D., Saluzzi, M., Appenzeller, S., Rittner, L., Frayne, R., Lotofo, R., 2018. An open, multi-vendor, multi-field-strength brain MR dataset and analysis of publicly available skull stripping methods agreement. NeuroImage 170, 482–494. d...

  209. [223]

    Digital Retinal Image for Vessel Extraction (DRIVE) Database

    Staal, J., Abr` amoff, M., Niemeijer, M., Viergever, M., Ginneken, B., 2013. Digital Retinal Image for Vessel Extraction (DRIVE) Database

  210. [224]

    Sun, Y., Wang, X., Tang, X., 2014. Deep learning face representation from predicting 10,000 classes, in: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE Computer Society. pp. 1891–1898. doi: 10.1109/CVPR.2014.244

  211. [225]

    Learning and Example Selection for Object and Pattern Detection

    Sung, K.k., 1996. Learning and Example Selection for Object and Pattern Detection. PhD thesis , 195doi: https: //doi.org/10.1016/j.comnet.2014.12.002

  212. [226]

    DeepID3: Face Recognition with Very Deep Neural Networks URL: http://arxiv.org/abs/1502.00873

    Sun, Y., Liang, D., Wang, X., Tang, X., 2015. DeepID3: Face Recognition with Very Deep Neural Networks URL: http://arxiv.org/abs/1502.00873

  213. [227]

    Deep Learning Face Representation by Joint Identification-Verification

    Sun, Y., Wang, X., Tang, X., . Deep Learning Face Representation by Joint Identification-Verification. Technical Report

  214. [228]

    DeepFace: Closing the Gap to Human-Level Performance in Face Verification

    Taigman, Y., Marc’, M.Y., Ranzato, A., Wolf, L., . DeepFace: Closing the Gap to Human-Level Performance in Face Verification. Technical Report

  215. [229]

    Face recognition: Past, present and future (a review)

    Taskiran, M., Kahraman, N., Erdem, C.E., 2020. Face recognition: Past, present and future (a review). Digital Signal Processing 106, 102809. URL: https://www.sciencedirect.com/science/article/pii/S1051200420301548, doi:https: //doi.org/10.1016/j.dsp.2020.102809

  216. [230]

    Snapshot Serengeti, high-frequency annotated camera trap images of 40 mammalian species in an African savanna

    Swanson, A., Kosmala, M., Lintott, C., Simpson, R., Smith, A., Packer, C., 2015. Snapshot Serengeti, high-frequency annotated camera trap images of 40 mammalian species in an African savanna. Scientific Data 2, 1–14. doi: 10.1038/ sdata.2015.26

  217. [231]

    Deep semantic segmentation of natural and medical images: A review

    Taghanaki, S.A., Abhishek, K., Cohen, J.P., Cohen-Adad, J., Hamarneh, G., 2020. Deep semantic segmentation of natural and medical images: A review

  218. [232]

    Superparsing: Scalable nonparametric image parsing with superpixels

    Tighe, J., Lazebnik, S., 2013. Superparsing: Scalable nonparametric image parsing with superpixels. International Journal of Computer Vision 101, 329–349. doi: 10.1007/s11263-012-0574-z

  219. [233]

    80 million tiny images: A large data set for nonparametric object and scene recognition

    Torralba, A., Fergus, R., Freeman, W.T., 2008. 80 million tiny images: A large data set for nonparametric object and scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 30, 1958–1970. doi: 10.1109/TPAMI. 2008.128

  220. [234]

    YFCC100M: The 30 new data in multimedia research

    Thomee, B., Elizalde, B., Shamma, D.A., Ni, K., Friedland, G., Poland, D., Borth, D., Li, L.J., 2016. YFCC100M: The 30 new data in multimedia research. Communications of the ACM 59, 64–73. doi: 10.1145/2812802

  221. [235]

    SuperParsing: Scalable Nonparametric Image Parsing with Superpixels, in: Daniilidis, K., Maragos, P., Paragios, N

    Tighe, J., Lazebnik, S., 2010. SuperParsing: Scalable Nonparametric Image Parsing with Superpixels, in: Daniilidis, K., Maragos, P., Paragios, N. (Eds.), Computer Vision – ECCV 2010, Springer Berlin Heidelberg, Berlin, Heidelberg. pp. 352–365

  222. [236]

    EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos

    Twinanda, A.P., Shehata, S., Mutter, D., Marescaux, J., De Mathelin, M., Padoy, N., 2017. EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos. IEEE Transactions on Medical Imaging 36, 86–97. doi:10.1109/TMI.2016.2593957

  223. [237]

    KiTS19 Challenge

    University of Minnesota, University of Melbourne, 2019. KiTS19 Challenge. URL: https://kits19.grand-challenge. org/

  224. [238]

    Sharing features: Efficient boosting procedures for multiclass object detection

    Torralba, A., Murphy, K.P., Freeman, W.T., 2004. Sharing features: Efficient boosting procedures for multiclass object detection. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2. doi:10.1109/cvpr.2004.1315241

  225. [239]

    Data descriptor: The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions

    Tschandl, P., Rosendahl, C., Kittler, H., 2018. Data descriptor: The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Scientific Data 5, 1–9. doi: 10.1038/sdata.2018.161

  226. [240]

    Robust Real-time Object Detection

    Viola, P., Viola, P., Jones, M., 2001b. Robust Real-time Object Detection. INTERNATIONAL JOURNAL OF COM- PUTER VISION

  227. [241]

    The Caltech-ucsd Birds-200-2011 Dataset

    Wah, C., Branson, S., Welinder, P., Perona, P., Belongie, S., 2011. The Caltech-ucsd Birds-200-2011 Dataset

  228. [242]

    Autonomous vehicle perception: The technology of today and tomorrow

    Van Brummelen, J., O’Brien, M., Gruyer, D., Najjaran, H., 2018. Autonomous vehicle perception: The technology of today and tomorrow. Transportation Research Part C: Emerging Technologies 89, 384–406. URL: https://doi.org/10. 1016/j.trc.2018.02.012, doi:10.1016/j.trc.2018.02.012

  229. [243]

    Rapid object detection using a boosted cascade of simple features

    Viola, P., Viola, P., Jones, M., 2001a. Rapid object detection using a boosted cascade of simple features. ACCEPTED CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION 2001

  230. [244]

    TorontoCity: Seeing the World with a Million Eyes

    Wang, S., Bai, M., Mattyus, G., Chu, H., Luo, W., Yang, B., Liang, J., Cheverie, J., Fidler, S., Urtasun, R., . TorontoCity: Seeing the World with a Million Eyes. Technical Report

  231. [245]

    TorontoCity : Seeing the World with a Million Eyes

    Wang, S., Bai, M., Mattyus, G., Chu, H., Luo, W., Yang, B., Liang, J., Cheverie, J., Fidler, S., Urtasun, R., 2016. TorontoCity : Seeing the World with a Million Eyes

  232. [246]

    Learning to detect salient objects with image-level supervision , 3796–3805doi: 10.1109/CVPR.2017.404

    Wang, L., Lu, H., Wang, Y., Feng, M., Wang, D., Yin, B., Ruan, X., 2017a. Learning to detect salient objects with image-level supervision , 3796–3805doi: 10.1109/CVPR.2017.404

  233. [247]

    The ApolloScape Open Dataset for Autonomous Driving and its Application

    Wang, P., Huang, X., Cheng, X., Zhou, D., Geng, Q., Yang, R., 2019. The ApolloScape Open Dataset for Autonomous Driving and its Application. IEEE Transactions on Pattern Analysis and Machine Intelligence , 1–1doi: 10.1109/tpami. 2019.2926463

  234. [248]

    Cancer Digital Slide Archive

    Winship Cancer Institute, . Cancer Digital Slide Archive. URL: https://cancer.digitalslidearchive.org/

  235. [249]

    Face recognition in unconstrained videos with matched background similarity

    Wolf, L., Hassner, T., Maoz, I., 2011. Face recognition in unconstrained videos with matched background similarity. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , 529–534doi: 10. 1109/CVPR.2011.5995566

  236. [250]

    ChestX-ray8: Hospital-scale chest X-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases

    Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., Summers, R.M., 2017b. ChestX-ray8: Hospital-scale chest X-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. Proceedings - 30th IEEE Conference on Computer Vision and Patt...

  237. [251]

    IARPA Janus Benchmark-B Face Dataset

    Whitelam, C., Taborsky, E., Blanton, A., Maze, B., Adams, J., Miller, T., Kalka, N., Jain, A.K., Duncan, J.A., Allen, K., Cheney, J., Grother, P., 2017. IARPA Janus Benchmark-B Face Dataset. IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops ...

  238. [252]

    IP102: A large-scale benchmark dataset for insect pest recognition, in: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp

    Wu, X., Zhan, C., Lai, Y.K., Cheng, M.M., Yang, J., 2019. IP102: A large-scale benchmark dataset for insect pest recognition, in: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 8779–8788. doi: 10.1109/CVPR.2019.00899

  239. [253]

    What is and what is not a salient object? learning salient object detector by ensembling linear exemplar regressors , 4399–4407doi: 10.1109/CVPR.2017.468

    Xia, C., Li, J., Chen, X., Zheng, A., Zhang, Y., 2017a. What is and what is not a salient object? learning salient object detector by ensembling linear exemplar regressors , 4399–4407doi: 10.1109/CVPR.2017.468

  240. [254]

    Synscapes: A Photorealistic Synthetic Dataset for Street Scene Parsing URL: http: //arxiv.org/abs/1810.08705

    Wrenninge, M., Unger, J., 2018. Synscapes: A Photorealistic Synthetic Dataset for Street Scene Parsing URL: http: //arxiv.org/abs/1810.08705

  241. [255]

    Automatic Landmark Estimation for Adolescent Idiopathic Scoliosis Assessment Using BoostNet, in: Medical Image Computing and Computer Assisted Intervention - MICCAI, pp

    Wu, H., Bailey, C., Rasoulinejad, P., Li, S., 2017. Automatic Landmark Estimation for Adolescent Idiopathic Scoliosis Assessment Using BoostNet, in: Medical Image Computing and Computer Assisted Intervention - MICCAI, pp. 127–135

  242. [256]

    SUN database: Large-scale scene recognition from abbey to zoo

    Xiao, J., Hays, J., Ehinger, K.A., Oliva, A., Torralba, A., 2010. SUN database: Large-scale scene recognition from abbey to zoo. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , 3485–3492doi:10.1109/CVPR.2010.5539970. 31

  243. [257]

    REtroSpective Evaluation of Cerebral Tumors (RESECT): A clinical database of pre-operative MRI and intra-operative ultrasound in low-grade glioma surgeries: A

    Xiao, Y., Fortin, M., Unsg¨ ard, G., Rivaz, H., Reinertsen, I., 2017. REtroSpective Evaluation of Cerebral Tumors (RESECT): A clinical database of pre-operative MRI and intra-operative ultrasound in low-grade glioma surgeries: A. Medical Physics 44, 3875–3882. doi: 10.1002/mp.12268

  244. [258]

    DOTA: A Large-scale Dataset for Object Detection in Aerial Images

    Xia, G.S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., Zhang, L., 2017b. DOTA: A Large-scale Dataset for Object Detection in Aerial Images. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , 3974–3983...

  245. [259]

    AID: A Benchmark Dataset for Performance Evaluation of Aerial Scene Classification

    Xia, G.S., Hu, J., Hu, F., Shi, B., Bai, X., Zhong, Y., Zhang, L., 2016. AID: A Benchmark Dataset for Performance Evaluation of Aerial Scene Classification. IEEE Transactions on Geoscience and Remote Sensing 55, 3965–3981. URL: http://arxiv.org/abs/1608.05167http://dx.doi.org/...

  246. [260]

    YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark , 1–10URL: http://arxiv.org/abs/1809.03327

    Xu, N., Yang, L., Fan, Y., Yue, D., Liang, Y., Yang, J., Huang, T., 2018b. YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark , 1–10URL: http://arxiv.org/abs/1809.03327

  247. [261]

    Hierarchical saliency detection on extended cssd

    Yan, Q., Shi, J., Xu, L., Jia, J., 2014. Hierarchical saliency detection on extended cssd. IEEE Transactions on Pattern Analysis and Machine Intelligence 38. doi: 10.1109/TPAMI.2015.2465960

  248. [262]

    PointFusion: Deep Sensor Fusion for 3D Bounding Box Estimation

    Xu, D., Anguelov, D., Jain, A., 2018a. PointFusion: Deep Sensor Fusion for 3D Bounding Box Estimation. Technical Report

  249. [263]

    CAMEL: A weakly supervised learn- ing framework for histopathology image segmentation

    Xu, G., Song, Z., Sun, Z., Ku, C., Yang, Z., Liu, C., Wang, S., Ma, J., Xu, W., 2019. CAMEL: A weakly supervised learn- ing framework for histopathology image segmentation. Proceedings of the IEEE International Conference on Computer Vision 2019-Octob, 10681–10690. doi: 10.110...

  250. [264]

    Learning Face Representation from Scratch URL: http://arxiv.org/abs/1411

    Yi, D., Lei, Z., Liao, S., Li, S.Z., 2014. Learning Face Representation from Scratch URL: http://arxiv.org/abs/1411. 7923

  251. [265]

    BDD100K: A Diverse Driving Video Database with Scalable Annotation Tooling , 1–16

    Yu, F., Xian, W., Chen, Y., Liu, F., Liao, M., Madhavan, V., Darrell, T., 2018. BDD100K: A Diverse Driving Video Database with Scalable Annotation Tooling , 1–16

  252. [266]

    Saliency detection via graph-based manifold ranking

    Yang, C., Zhang, L., Lu, H., Ruan, X., Yang, M.H., 2013. Saliency detection via graph-based manifold ranking. Pro- ceedings / CVPR, IEEE Computer Society Conference on Computer Vision and Pattern Recognition. IEEE Computer Society Conference on Computer Vision and Pattern Reco...

  253. [267]

    A multi-center milestone study of clinical vertebral CT segmentation

    Yao, J., Burns, J.E., Forsberg, D., Seitel, A., Rasoulian, A., Abolmaesumi, P., Hammernik, K., Urschler, M., Ibragimov, B., Korez, R., Vrtovec, T., Castro-Mateos, I., Pozo, J.M., Frangi, A.F., Summers, R.M., Li, S., 2016. A multi-center milestone study of clinical vertebral CT...

  254. [268]

    Salient object subitizing , 4045–4054doi: 10.1109/CVPR.2015.7299031

    Zhang, J., Ma, S., Sameki, M., Sclaroff, S., Betke, M., Lin, Z., Shen, X., Price, B., Mech, R., 2015. Salient object subitizing , 4045–4054doi: 10.1109/CVPR.2015.7299031

  255. [269]

    Capsal: Leveraging captioning to boost semantics for salient object detection , 6017–6026doi: 10.1109/CVPR.2019.00618

    Zhang, L., Zhang, J., Lin, Z., Lu, H., He, Y., 2019. Capsal: Leveraging captioning to boost semantics for salient object detection , 6017–6026doi: 10.1109/CVPR.2019.00618

  256. [270]

    Mutual graph learning for camouflaged object detection

    Zhai, Q., Li, X., Yang, F., Chen, C., Cheng, H., Fan, D.P., 2021. Mutual graph learning for camouflaged object detection

  257. [271]

    INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps

    Zhan, W., Sun, L., Wang, D., Shi, H., Clausse, A., Naumann, M., Kummerle, J., Konigshof, H., Stiller, C., de La Fortelle, A., Tomizuka, M., 2019. INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps

  258. [272]

    ModaNet: A large-scale street fashion dataset with polygon annotations

    Zheng, S., Hadi Kiapour, M., Yang, F., Piramuthu, R., 2018. ModaNet: A large-scale street fashion dataset with polygon annotations. MM 2018 - Proceedings of the 2018 ACM Multimedia Conference , 1670–1678doi:10.1145/3240508.3240652

  259. [273]

    Places: An Image Database for Deep Scene Understanding

    Zhou, B., Lapedriza, A., Torralba, A., Oliva, A., 2017a. Places: An Image Database for Deep Scene Understanding. Journal of Vision 17, 296. doi: 10.1167/17.10.296

  260. [274]

    CityPersons: A Diverse Dataset for Pedestrian Detection

    Zhang, S., Benenson, R., Schiele, B., 2017. CityPersons: A Diverse Dataset for Pedestrian Detection. Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 2017-January, 4457–4465. URL: http://arxiv.org/abs/1702.05693

  261. [275]

    Object detection with deep learning: A review

    Zhao, Z.Q., Zheng, P., Xu, S.T., Wu, X., 2019. Object detection with deep learning: A review. IEEE Transactions on Neural Networks and Learning Systems 30, 3212–3232. doi: 10.1109/TNNLS.2018.2876865

  262. [276]

    Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not? Technical Report

    Zhou, E., Yin, Q., . Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not? Technical Report

  263. [277]

    Zhu, H., Chen, X., Dai, W., Fu, K., Ye, Q., Jiao, J., 2015. Orientation robust object detection in aerial images using deep convolutional neural network, in: Proceedings - International Conference on Image Processing, ICIP, IEEE Computer Society. pp. 3735–3739. doi: 10.1109/IC...

  264. [278]

    Scene parsing through ADE20K dataset

    Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A., 2017b. Scene parsing through ADE20K dataset. Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 2017-Janua, 5122–5130. doi:10.1109/CVPR.2017.544

  265. [279]

    Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not? URL: http://arxiv.org/abs/1501.04690

    Zhou, E., Cao, Z., Yin, Q., 2015. Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not? URL: http://arxiv.org/abs/1501.04690

  266. [280]

    Object Detection in 20 Years: A Survey , 1–39URL: http://arxiv.org/abs/ 1905.05055

    Zou, Z., Shi, Z., Guo, Y., Ye, J., 2019b. Object Detection in 20 Years: A Survey , 1–39URL: http://arxiv.org/abs/ 1905.05055. 32

  267. [282]

    FashionAI: A Hierarchical Dataset for Fashion Understanding

    Zou, X., Kong, X., Wong, W., Wang, C., Liu, Y., Cao, Y., 2019a. FashionAI: A Hierarchical Dataset for Fashion Understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops

  268. [283]

    Random access memories: A new paradigm for target detection in high resolution aerial remote sensing images

    Zou, Z., Shi, Z., 2018. Random access memories: A new paradigm for target detection in high resolution aerial remote sensing images. IEEE Transactions on Image Processing 27, 1100–1111. doi: 10.1109/TIP.2017.2773199

  269. [2015]

    Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 07-12-June, 1931–1939

    Pushing the frontiers of unconstrained face detection and recognition: IARPA Janus Benchmark A. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 07-12-June, 1931–1939. doi: 10. 1109/CVPR.2015.7298803

  270. [2019]

    nuScenes: A multimodal dataset for autonomous driving

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.