Pith. sign in

REVIEW 3 major objections 4 minor 92 references

Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper introduces ODOR, a dataset of 38,116 object-level annotations across 4,712 artworks and 139 fine-grained categories, built to make object detection face small, dense, off-center, and smell-related objects in paintings.

desk verdict A genuinely useful new benchmark for artwork object detection, but its fine-grained label quality is unmeasured and needs to be demonstrated before the difficulty claims are taken at face value. read the letter →

arxiv 2507.08384 v1 pith:3RNN5QZE submitted 2025-07-11 cs.CV

classification cs.CV MSC 68T45 PACS 42.30.Tz42.30.Sy
keywords ObjectDetectionDatasetComputationalHumanitiesOlfactionSmallTransferLearningArtworkFine-grainedClassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ODOR is a new object-detection dataset for artworks, built around smell-related imagery: 38,116 object-level annotations across 4,712 images and 139 fine-grained categories such as flower species, fruits, drinking vessels, smoking equipment, and smoke-related phenomena. The paper's central claim is that ODOR is the second-largest public artwork dataset with object-level annotations by image count and the most complex by class count and annotation density, with objects spread across the full canvas rather than concentrated in the centre. To back this up, the authors compare its statistics to existing artwork datasets and run baseline object detectors, reporting that the strongest model reaches only 22.6 AP, while evaluating at supercategory level improves scores by 35 to 60 percent. The intended payoff is a reusable benchmark that pushes detectors toward small, occluded, peripheral objects and supports quantitative research in olfactory cultural heritage.

What carries the argument

The load-bearing mechanism is the dataset's two-level annotation hierarchy combined with its acquisition and analysis protocol. All 139 fine-grained classes sit under supercategories chosen for visual and olfactory similarity rather than biological taxonomy, so a detector can fall back to a coarse prediction when species-level confidence is low; this hierarchy is what makes the supercategory-evaluation experiment meaningful. The authors also use standard COCO-format annotations, a crowd-annotation pipeline with multiple rounds of expert art-history correction, and statistical characterizations of spatial distribution, occlusion, and box-size histograms to establish the dataset's difficulty. On top of this, a benchmark of five detector families (two-stage, transformer-based, and one-stage) with varying backbones provides the performance numbers that make the challenge concrete.

What would settle it

Independently re-annotate a random sample of about 200 ODOR test images with two art-history experts and measure per-class agreement on the fine-grained categories; low agreement on classes such as flower species or drinking vessels would mean the class-level AP baselines partly measure annotation noise rather than detector ability. In parallel, a chi-square goodness-of-fit test comparing object-centre positions to a uniform distribution over the canvas would settle whether the claimed full-canvas spread is real or an artifact of visual inspection.

Watch

Extended reading notes

Core claim

The central discovery is the dataset itself. ODOR contains 38,116 object-level bounding-box annotations in 4,712 historical artworks, organized into 139 fine-grained classes under eight pragmatic supercategories (flower, fruit, vegetable, mammal, drinking vessel, smoking equipment, smoke-related, and other). Its distinctive statistical properties are dense and overlapping instances (on average 8.1 boxes and 3.7 classes per image), a long-tailed class distribution in which some classes have only 16 training instances, a high share of very small objects covering less than 4 percent of image area, and object centres spread across the whole canvas instead of clustered in the middle. The paper demonstrates that these properties make detection genuinely hard: the best evaluated configuration, a DINO detector with a FocalNet-L backbone, reaches 22.6 AP; relaxing the task to supercategories raises AP substantially, showing that fine-grained distinctions are a main source of difficulty. The authors position ODOR as the second-largest public artwork detection dataset by image count and the most complex in terms of class count and annotation density.

Load-bearing premise

The whole benchmark rests on the accuracy of the 38,116 fine-grained labels, which were crowd-annotated and expert-corrected but with no reported inter-annotator agreement, spot-check error rate, or formal QA protocol, so every AP number inherits any unmeasured label noise.

Editorial extensions

If this is right

  • The best baseline reaching only 22.6 AP shows that current detectors leave a wide margin on dense, fine-grained, small-object artwork detection.
  • The 35 to 60 percent gain from evaluating at supercategory level means fine-grained class boundaries, not localisation, are the dominant source of difficulty for the strongest models.
  • Detection-specific pretraining on natural images appears to transfer well: a COCO-pretrained one-stage detector roughly matches much heavier transformer backbones, especially on small objects.
  • The dataset enables quantitative art-history studies of olfactory references, such as tracing the prevalence of roses in still-life paintings over centuries.
  • Researchers needing robustness to off-centre, occluded, and small objects can use ODOR as a complement to natural-image benchmarks such as COCO.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because label noise is unreported, a practical next step is measuring inter-annotator agreement on the fine-grained classes; if it is low, category-level AP rankings may need to be re-read as rankings of annotation ambiguity.
  • The supercategory hierarchy invites hierarchical or open-vocabulary detectors: a model that predicts the coarse class with calibrated uncertainty could be more useful for art historians than one forced to guess among 139 species-level labels.
  • The dataset's class distribution reflects olfactory relevance, not general art-historical salience, so detectors trained on it may not transfer to arbitrary artwork-detection tasks without reweighting or fine-tuning.
  • Because several images are only available through external source links, benchmark reproducibility depends on link persistence; hashed copies or a complete mirror would make future comparisons more robust.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces ODOR, a dataset of 4,712 artwork images with 38,116 object-level bounding-box annotations across 139 fine-grained categories and 8 supercategories, collected in the Odeuropa project with a focus on olfactory references. The authors report statistical analyses of category distribution, spatial spread, occlusion, and object sizes, and provide extensive baselines for five detector families (F-RCNN, DINO, YOLO-v8, FCOS, MADet), including supercategory-wise and class-wise evaluations. The paper claims that ODOR is the second-largest public artwork detection dataset by image count and the most complex by class count and annotation density, and that it offers a challenging benchmark for small, occluded, and peripheral objects in artworks.

Significance. If the label-quality concerns are resolved, ODOR would be a valuable and distinctive public resource for the computational humanities: it is the only artwork detection dataset focused on fine-grained olfactory-relevant categories, it exhibits a genuinely long-tailed distribution and dense, overlapping instances, and it is released with metadata, download tools, and reproducible baseline code. The baseline study is a concrete strength: the authors evaluate multiple modern detectors, report size- and supercategory-stratified AP, and make code and data publicly available (Zenodo, GitHub, HuggingFace), which lowers the barrier for follow-up work. The dataset has already been used in the ICPR 2022 ODOR challenge, indicating real community uptake. The central claims about dataset complexity and difficulty are, however, currently weakened by the absence of quantitative label-quality evidence and by at least one inaccurate comparative statement.

major comments (3)
  1. [§3.2, §5] The annotation process is described only qualitatively: Section 3.2 states that AMT annotations were 'manually checked by experts' with 'multiple rounds of corrections', and Section 5 concedes that 'an inherent level of ambiguity persists.' No inter-annotator agreement, spot-check error rate, per-class confusion statistics, or adjudication protocol is reported. This is load-bearing because the dataset's headline properties—139 fine-grained categories, subtle distinctions among flower species and drinking vessels—and all downstream quantities (baseline AP in Tables 4–6, class-wise AP and Spearman correlations in Table 7, and the supercategory gains in Table 6) inherit the unmeasured label noise. I request a quantitative QA component, such as Cohen's kappa on a held-out subset, a per-class error analysis on a random sample, or a sensitivity analysis showing that the main conclusions are robust to plausible label-error rates.
  2. [§6, Table 1] The conclusion states that ODOR is 'the second-largest public artwork dataset with object-level annotations in terms of the number of images.' This is contradicted by the paper's own Table 1, which lists DEArt with 15,157 images and Human-Art with 50,000 images—both larger than ODOR's 4,712. The claim must be corrected or explicitly qualified (e.g., excluding person-only or semi-automatically annotated datasets), otherwise the central positioning of the dataset is overstated.
  3. [§3.4, Fig. 7] The 'spreaded' property—one of the three characteristics advertised in the title—is supported only by visual comparison of heatmaps in Fig. 7. A quantitative test comparing the normalized object-center distributions of ODOR with PeopleArt, IconArt, and DEArt (e.g., a chi-square or Kolmogorov–Smirnov test) would make the claim falsifiable and is necessary to substantiate the full-canvas distribution assertion, which the paper explicitly contrasts with the center-biased distributions of other datasets.
minor comments (4)
  1. [Title] The title has a spacing typo: 'TheObject Detection' should be 'The Object Detection.'
  2. [§3.1] The sentence 'On average they have a width of 653 and a width of 641 pixels' should presumably read 'a width of 653 and a height of 641 pixels.'
  3. [§4.1] The phrase 'Both FCOS is and YOLO-v8 are anchor-free' should be 'Both FCOS and YOLO-v8 are anchor-free.'
  4. [Table 7] Several rows in Table 7 contain truncated numeric values with ellipses (e.g., '7204...' and '7086...'); these should be formatted as complete numbers or the table should state that truncated values are rounded for display.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a data/benchmark contribution whose statistics and baseline results are independent measurements, not derivations from their own outputs.

full rationale

The paper is a dataset and benchmark contribution, not a derivation chain in which an output quantity is defined in terms of the quantity it predicts. The central statistics (38,116 instances, 139 classes, 4,712 images, density and spatial-distribution properties) are direct measurements of the released annotations and of the images, so there is no self-definitional reduction. The baseline results in Tables 4-6 and the class-wise analysis in Table 7 are empirical evaluations of standard detectors on a fixed train/test split; they involve no fitted parameter that is afterwards renamed as a prediction, and the test AP numbers are not used to construct the dataset statistics. The supercategory evaluation in Table 6 is a different scoring protocol applied to the same detector outputs, not a quantity that was fitted to produce those outputs. Self-citations to the earlier ODOR challenge [10] and to SniffyArt [7] are contextual references to prior project versions; the central claim that ODOR is a large, dense, fine-grained artwork detection benchmark does not depend on those citations because the dataset itself is released and the baseline numbers are reproducible measurements. The acknowledged label ambiguity in Section 5 ('an inherent level of ambiguity persists') is a quality and validity limitation, but it is not a circularity: no claim reduces by construction to the annotation process, and the absence of inter-annotator agreement statistics is a correctness/evidence concern, not a circular-reasoning concern. The comparison to PeopleArt, IconArt, DEArt, PoPArt, Human-Art, and SniffyArt is an external empirical comparison based on reported dataset statistics, not a self-referential derivation. Therefore the appropriate finding is no significant circularity.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

Dataset papers shift the ledger from fitted constants to curation choices. The two numeric hand-picks (16-instance cutoff, 90:10 split) plus the disclosed supercategory grouping define the released benchmark and hence all reported AP numbers; the domain assumptions about smell-relevance and label accuracy are the real epistemic load-bearing items. No new physical or conceptual entities are postulated.

free parameters (3)
  • minimum instance threshold = 16 instances across at least 3 images
    Hand-chosen cutoff in Section 3.5 that decides which of 227 collected categories survive into the released 139-category benchmark; it directly shapes the long-tailed class distribution presented as a core challenge.
  • train/test split ratio = 90:10
    Hand-chosen split in Section 3.5 (4264 train / 448 test images); all baseline numbers depend on this split definition.
  • supercategory grouping rule = pragmatic visual similarity, e.g., whales grouped as fish-like animals
    Disclosed ad hoc hierarchy (Section 3.2) that determines the supercategory evaluation in Table 6 and the reported AP increases; a hand-made grouping decision with direct effect on reported metrics.
assumptions (4)
  • domain assumption The Odeuropa controlled vocabularies [59] operationalize 'olfactory references' in art, and the smell-related query terms used for collection define a valid target distribution for detection benchmarks.
    Section 3.2: the semantic framing of the entire dataset depends on this premise; the collection strategy is purposive, not a random sample of artworks.
  • domain assumption A fixed two-level taxonomy can meaningfully categorize historical objects despite artistic abstraction, historical change, and borderline cases.
    Section 5 self-discloses that artists can 'deliberately introduc[e] ambiguity or uncertainty' and that an 'inherent level of ambiguity persists', yet every instance is assigned to a rigid class.
  • domain assumption Expert review of crowd annotations yields accurate labels for 38,116 boxes across 139 fine-grained classes.
    Section 3.2 asserts manual expert checking and 'multiple rounds of corrections' but provides no inter-annotator agreement, error rates, or QA protocol.
  • standard math The COCO evaluation protocol (AP over IoU 0.50-0.95, COCO size bins) is the appropriate metric for quantifying benchmark difficulty.
    Sections 4.1 and 4.2: standard practice that makes baseline numbers comparable to the broader detection literature; a reasonable choice rather than the paper's own new claim.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset." pith.science (2026). https://pith.science/paper/3RNN5QZE

@misc{pith2026250708384,
  author       = {Pith},
  title        = {Pith review of: Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3RNN5QZE}},
  note         = {Machine review of arXiv:2507.08384}
}
read the original abstract

Real-world applications of computer vision in the humanities require algorithms to be robust against artistic abstraction, peripheral objects, and subtle differences between fine-grained target classes. Existing datasets provide instance-level annotations on artworks but are generally biased towards the image centre and limited with regard to detailed object classes. The proposed ODOR dataset fills this gap, offering 38,116 object-level annotations across 4712 images, spanning an extensive set of 139 fine-grained categories. Conducting a statistical analysis, we showcase challenging dataset properties, such as a detailed set of categories, dense and overlapping objects, and spatial distribution over the whole image canvas. Furthermore, we provide an extensive baseline analysis for object detection models and highlight the challenging properties of the dataset through a set of secondary studies. Inspiring further research on artwork object detection and broader visual cultural heritage studies, the dataset challenges researchers to explore the intersection of object recognition and smell perception.

Figures

Figures reproduced from arXiv: 2507.08384 by the authors.

Figure 1
Figure 1. Examples of the ODOR dataset showing its variety with dense and overlapping [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Distribution of image widths (blue, dashed) and heights (red) of the artworks in the [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Class distribution plot for the ODOR dataset train (blue) and test (red) splits; the [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Artwork from the ODOR dataset with rich image description. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Exemplary metadata for an artwork from the dataset. For brevity, only selected [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Number of annotated categories (top), and instances (bottom) per image for DEArt, [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Spatial distribution of object centres in normalised image coordinates of various [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Percentage of instances per number of overlapping instances (top), and area of [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Examples from the DEArt [6] dataset (left) and the ODOR dataset (right) displaying varying amounts of occlusion. The instances of the eliminated categories are either removed from the dataset or assigned to their respective supercategory, depending on whether the class…
Figure 10
Figure 10. Figure 10: Histogram of instance box sizes, relative to the image size. [PITH_FULL_IMAGE:figures/full_fig_p016_10.png]
Figure 11
Figure 11. Figure 11: Example detections on a heap of small, partly occluding apples. Zoomed in for [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]
Figure 12
Figure 12. Figure 12: a A bouquet of various flower species; b–d different model predictions. For [PITH_FULL_IMAGE:figures/full_fig_p021_12.png]
Figure 13
Figure 13. Figure 13: Jan Steen Jug (left) and occluding bunches of grape, partly blending into the [PITH_FULL_IMAGE:figures/full_fig_p023_13.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

92 extracted references · 72 canonical work pages

  1. [1]

    H. Cai, Q. Wu, T. Corradi, P. Hall, The Cross-Depiction Problem: Computer Vision Algorithms for Recognising Objects in Artwork and in Photographs, arXiv preprint arXiv:1505.00110 (2015)

  2. [2]

    Madhu, A

    P. Madhu, A. Villar-Corrales, R. Kosti, T. Bendschus, C. Reinhardt, P. Bell, A. Maier, V. Christlein, Enhancing Human Pose Estimation in Ancient Vase Paintings via Perceptually-grounded Style Transfer Learning, ACM Journal on Computing and Cultural Heritage 16 (1) (2022) 1–17

  3. [3]

    Strezoski, M

    G. Strezoski, M. Worring, OmniArt: A Large-scale Artistic Benchmark, ACM Transactions on Multimedia Computing, Communications, and Ap- plications (TOMM) 14 (4) (2018) 1–21

  4. [4]

    Westlake, H

    N. Westlake, H. Cai, P. Hall, Detecting People in Artwork with CNNs, in: Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part I 14, Springer, 2016, pp. 825–841

  5. [5]

    Gonthier, Y

    N. Gonthier, Y. Gousseau, S. Ladjal, O. Bonfait, Weakly Supervised Object Detection in Artworks, in: L. Leal-Taix´ e, S. Roth (Eds.), Computer Vision – ECCV 2018 Workshops, Springer International Publishing, Cham, 2019, pp. 692–709

  6. [6]

    Reshetnikov, M.-C

    A. Reshetnikov, M.-C. Marinescu, J. M. Lopez, DEArt: Dataset of European Art, in: European Conference on Computer Vision, Springer, 2022, pp. 218– 233

  7. [7]

    Zinnen, A

    M. Zinnen, A. Hussian, H. Tran, P. Madhu, A. Maier, V. Christlein, Sniff- yArt: The Dataset of Smelling Persons, in: Proceedings of the 5th Workshop on AnalySis, Understanding and ProMotion of HeritAge Contents, SUMAC ’23, Association for Computing Machinery, New York, NY, USA, 2023, p. 49–58. doi:10.1145/3607542.3617357. URL https://doi.org/10.1145/36075...

  8. [8]

    X. Ju, A. Zeng, J. Wang, Q. Xu, L. Zhang, Human-Art: A versatile human- centric dataset bridging natural and artificial scenes, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 618–629

Show all 92 references
  1. [9]

    Gupta, P

    A. Gupta, P. Dollar, R. Girshick, L VIS: A Dataset for Large Vocabulary Instance Segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 5356–5364

  2. [10]

    Zinnen, P

    M. Zinnen, P. Madhu, R. Kosti, P. Bell, A. Maier, V. Christlein, ODOR: The ICPR2022 Odeuropa Challenge on Olfactory Object Recognition, in: 2022 26th International Conference on Pattern Recognition (ICPR), IEEE, 2022, pp. 4989–4994. 26

  3. [11]

    S. Zhao, A. Akda˘ g Salah, A. A. Salah, Automatic Analysis of Human Body Representations in Western Art, in: Computer Vision–ECCV 2022 Workshops: Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part I, Springer, 2023, pp. 282–297

  4. [12]

    Kadish, S

    D. Kadish, S. Risi, A. S. Løvlie, Improving Object Detection in Art Images Using Only Style Transfer, in: 2021 International Joint Conference on Neural Networks (IJCNN), IEEE, 2021, pp. 1–8

  5. [13]

    P. Hall, H. Cai, Q. Wu, T. Corradi, Cross-depiction problem: Recognition and synthesis of photographs and artwork, Computational Visual Media 1 (2015) 91–103

  6. [14]

    Cetinic, J

    E. Cetinic, J. She, Understanding and Creating Art with AI: Review and Outlook, ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 18 (2) (2022) 1–22

  7. [15]

    Y. Lu, C. Guo, X. Dai, F.-Y. Wang, Data-efficient image captioning of fine art paintings via virtual-real semantic alignment training, Neurocomputing 490 (2022) 163–180

  8. [16]

    Sabatelli, M

    M. Sabatelli, M. Kestemont, W. Daelemans, P. Geurts, Deep Transfer Learning for Art Classification Problems, in: L. Leal-Taix´ e, S. Roth (Eds.), Computer Vision – ECCV 2018 Workshops, Springer International Publish- ing, Cham, 2019, pp. 631–646

  9. [17]

    Gonthier, Y

    N. Gonthier, Y. Gousseau, S. Ladjal, An analysis of the transfer learning of convolutional neural networks for artistic images, in: Pattern Recognition. ICPR International Workshops and Challenges: Virtual Event, January 10–15, 2021, Proceedings, Part III, Springer, 2021, pp. 546–561

  10. [18]

    Zinnen, P

    M. Zinnen, P. Madhu, P. Bell, A. Maier, V. Christlein, Transfer Learn- ing for Olfactory Object Detection, in: Digital Humanities Conference, 2022, Alliance of Digital Humanities Organizations, 2022, pp. 409–413, https://arxiv.org/abs/2301.09906

  11. [19]

    W. Zhao, W. Jiang, X. Qiu, Big Transfer Learning for Fine Art Classification, Computational Intelligence and Neuroscience 2022 (2022)

  12. [20]

    Cheng, J

    G. Cheng, J. Han, P. Zhou, D. Xu, Learning Rotation-Invariant and Fisher Discriminative Convolutional Neural Networks for Object Detection, IEEE Transactions on Image Processing 28 (1) (2019) 265–278. doi:10.1109/TI P.2018.2867198

  13. [21]

    Russakovsky, J

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al., Imagenet Large Scale Visual Recognition Challenge, International Journal of Computer Vision 115 (2015) 211–252. 27

  14. [22]

    T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ ar, C. L. Zitnick, Microsoft COCO: Common objects in context, in: Com- puter Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer, 2014...

  15. [23]

    Kuznetsova, H

    A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, et al., The Open Images Dataset v4: Unified image classification, object detection, and visual rela- tionship detection at scale, International Journal of ...

  16. [24]

    Cheng, X

    G. Cheng, X. Xie, J. Han, L. Guo, G.-S. Xia, Remote sensing image scene classification meets deep learning: Challenges, methods, benchmarks, and opportunities, IEEE Journal of Selected Topics in Applied Earth Observa- tions and Remote Sensing 13 (2020) 3735–3756

  17. [25]

    Cheng, C

    G. Cheng, C. Lang, M. Wu, X. Xie, X. Yao, J. Han, Feature Enhancement Network for Object Detection in Optical Remote Sensing Images, Journal of Remote Sensing (2021)

  18. [26]

    Crowley, A

    E. Crowley, A. Zisserman, The State of the Art: Object Retrieval in Paintings using Discriminative Regions, in: Proceedings of the British Machine Vision Conference. BMV A Press, 2014

  19. [27]

    E. J. Crowley, A. Zisserman, In Search of Art, in: Computer Vision- ECCV 2014 Workshops: Zurich, Switzerland, September 6-7 and 12, 2014, Proceedings, Part I 13, Springer, 2015, pp. 54–70

  20. [28]

    E. J. Crowley, A. Zisserman, The Art of Detection, in: Computer Vision– ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part I 14, Springer, 2016, pp. 721–737

  21. [29]

    Girshick, Fast R-CNN, in: Proceedings of the IEEE international confer- ence on computer vision, 2015, pp

    R. Girshick, Fast R-CNN, in: Proceedings of the IEEE international confer- ence on computer vision, 2015, pp. 1440–1448

  22. [30]

    Madhu, A

    P. Madhu, A. Meyer, M. Zinnen, L. M¨ uhrenberg, D. Suckow, T. Bendschus, C. Reinhardt, P. Bell, U. Verstegen, R. Kosti, et al., One-shot Object De- tection in Heterogeneous Artwork Datasets, in: 2022 Eleventh International Conference on Image Processing Theory, Tools and Appli...

  23. [31]

    Madhu, T

    P. Madhu, T. Marquart, R. Kosti, D. Suckow, P. Bell, A. Maier, V. Christlein, ICC++: Explainable feature learning for art history using image composi- tions, Pattern Recognition 136 (2023) 109153. doi:https://doi.org/10 .1016/j.patcog.2022.109153. URL https://www.sciencedirect...

  24. [32]

    Madhu, T

    P. Madhu, T. Marquart, R. Kosti, P. Bell, A. Maier, V. Christlein, Un- derstanding Compositional Structures in Art Historical Images Using Pose and Gaze Priors: Towards Scene Understanding in Digital Art History, in: Computer Vision–ECCV 2020 Workshops: Glasgow, UK, August 23–...

  25. [33]

    P. Bell, L. Impett, The Choreography of the Annunciation through a Computational Eye, Histoire de l’art 34 (87) (2021) 01–06

  26. [34]

    Bernasconi, GAB-Gestures for Artworks Browsing, in: 27th International Conference on Intelligent User Interfaces, 2022, pp

    V. Bernasconi, GAB-Gestures for Artworks Browsing, in: 27th International Conference on Intelligent User Interfaces, 2022, pp. 50–53

  27. [35]

    Impett, F

    L. Impett, F. Moretti, Totentanz. Operationalizing Aby Warburg’s Pathos- formeln, Tech. rep., Stanford Literary Lab (2017)

  28. [36]

    Impett, Analyzing Gesture in Digital Art History, in: The Routledge Companion to Digital Humanities and Art History, Routledge, 2020, pp

    L. Impett, Analyzing Gesture in Digital Art History, in: The Routledge Companion to Digital Humanities and Art History, Routledge, 2020, pp. 386–407

  29. [37]

    C. M. Becker, Aby Warburg’s Pathosformel as Methodological Paradigm, The Journal of Art Historiography 9 (2013) 9–CB1

  30. [38]

    S. Kim, J. Park, J. Bang, H. Lee, Seeing is Smelling: Localizing Odor- Related Objects in Images, in: Proceedings of the 9th Augmented Human International Conference, 2018, pp. 1–9

  31. [39]

    Y. Eda, H. Matsukura, Y. Nozaki, M. Sakamoto, Detection of odor-related objects in images based on everyday odors in Japan, in: Proceedings of the AAAI Spring Symposium: Socially Responsible AI for Well-being, 2023, pp. 59–60

  32. [40]

    Rodr ´ ıguez-Ortega, Image Processing and Computer Vision in the Field of Art History, in: The Routledge Companion to Digital Humanities and Art History, Routledge, 2020, pp

    N. Rodr ´ ıguez-Ortega, Image Processing and Computer Vision in the Field of Art History, in: The Routledge Companion to Digital Humanities and Art History, Routledge, 2020, pp. 338–357

  33. [41]

    N¨ aslund Dahlgren, A

    A. N¨ aslund Dahlgren, A. Wasielewski, Cultures of Digitization: A Historio- graphic Perspective on Digital Art History, Visual Resources 36 (4) (2020) 339–359

  34. [42]

    S. Lang, B. Ommer, Reflecting on How Artworks Are Processed and Ana- lyzed by Computer Vision, in: Proceedings of the European Conference on Computer Vision (ECCV) Workshops, 2018

  35. [43]

    Appadurai, The Social Life of Things, Tech

    A. Appadurai, The Social Life of Things, Tech. rep., Cambridge University Press (1988)

  36. [44]

    Hicks, M

    D. Hicks, M. C. Beaudry, The Oxford handbook of material culture studies, OUP Oxford, 2010. 29

  37. [45]

    M. J. Van Zuijlen, H. Lin, K. Bala, S. C. Pont, M. W. Wijntjes, Materials In Paintings (MIP): An interdisciplinary dataset for perception, art history, and computer vision, Plos one 16 (8) (2021) e0255109

  38. [46]

    Leemans, W

    I. Leemans, W. de Vries, Wind Trade: How the Concept of Wind Came to Embody Speculation in the Dutch Republic, The Journal of Modern History 94 (2) (2022) 288–325

  39. [47]

    Everingham, A

    M. Everingham, A. Zisserman, C. K. Williams, L. Van Gool, M. Allan, C. M. Bishop, O. Chapelle, N. Dalal, T. Deselaers, G. Dork´ o, et al., The 2005 PASCAL visual object classes challenge, in: Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classif...

  40. [48]

    Madhu, T

    P. Madhu, T. Marquart, R. Kosti, D. Suckow, P. Bell, A. Maier, V. Christlein, Icc++: Explainable feature learning for art history using image composi- tions, Pattern Recognition 136 (2023) 109153

  41. [49]

    Garcia, G

    N. Garcia, G. Vogiatzis, How to Read Paintings: Semantic Art Understand- ing with Multi-Modal Retrieval, in: Proceedings of the European Conference on Computer Vision (ECCV) Workshops, 2018

  42. [50]

    Garcia, C

    N. Garcia, C. Ye, Z. Liu, Q. Hu, M. Otani, C. Chu, Y. Nakashima, T. Mi- tamura, A Dataset and Baselines for Visual Question Answering on Art, in: Computer Vision–ECCV 2020 Workshops: Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16, Springer, 2020, pp. 92–108

  43. [51]

    M. J. Wilber, C. Fang, H. Jin, A. Hertzmann, J. Collomosse, S. Belongie, BAM! the Behance Artistic Media Dataset for Recognition Beyond Photog- raphy, in: Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 1202–1211

  44. [52]

    Wallace, D

    A. Wallace, D. McCarthy, Survey of GLAM Open Access Policy and Practice, https://docs.google.com/document/d/15U__Z50WCUM_OWQ9HKLvLMlk cMoCN68FLVl9OKJQ8yY/edit?usp=sharing, accessed: 2023-02-02 (2018)

  45. [53]

    Schneider, R

    S. Schneider, R. Vollmer, Poses of People in Art: A Data Set for Human Pose Estimation in Digital Art History, arXiv preprint arXiv:2301.05124 (2023)

  46. [54]

    Gonthier, IconArt Dataset (Oct

    N. Gonthier, IconArt Dataset (Oct. 2018). URL https://doi.org/10.5281/zenodo.4737435

  47. [55]

    Marinescu, A

    M.-C. Marinescu, A. Reshetnikov, J. M. L´ opez, Improving Object Detection in Paintings based on Time Contexts, in: 2020 International Conference on Data Mining Workshops (ICDMW), IEEE, 2020, pp. 926–932. 30

  48. [56]

    S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, J. Sun, Objects365: A Large-scale, High-quality Dataset for Object Detection, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 8430–8439

  49. [57]

    L. D. Couprie, Iconclass, a device for the iconographical analysis of art objects, Museum International 30 (3-4) (1978) 194–198

  50. [58]

    Brandhorst, E

    H. Brandhorst, E. Posthumus, Iconclass: a key to collaboration in the digital humanities, in: The Routledge Companion to Medieval Iconography, Routledge, 2016, pp. 201–218

  51. [59]

    Lisena, D

    P. Lisena, D. Schwabe, M. van Erp, R. Troncy, W. Tullett, I. Leemans, L. Marx, S. C. Ehrich, Capturing the Semantics of Smell: The Odeuropa Data Model for Olfactory Heritage Information, in: The Semantic Web: 19th International Conference, ESWC 2022, Hersonissos, Crete, Greece...

  52. [60]

    G. A. Miller, WordNet: a Lexical Database for English, Communications of the ACM 38 (11) (1995) 39–41

  53. [61]

    Redmon, A

    J. Redmon, A. Farhadi, YOLO9000: better, faster, stronger, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7263–7271

  54. [62]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sas- try, A. Askell, P. Mishkin, J. Clark, et al., Learning Transferable Visual Models from Natural Language Supervision, in: International Conference on Machine Learning, PMLR, 2021, pp. 8748–8763

  55. [63]

    Kamath, M

    A. Kamath, M. Singh, Y. LeCun, I. Misra, G. Synnaeve, N. Carion, MDETR – Modulated Detection for End-to-End Multi-Modal Understand- ing, arXiv:2104.12763 [cs]ArXiv: 2104.12763 (Apr. 2021). URL http://arxiv.org/abs/2104.12763

  56. [64]

    L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, K.-W. Chang, J. Gao, Grounded Language-Image Pre-training, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, New Orleans, LA, USA, 2022, pp. 109...

  57. [65]

    S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, L. Zhang, Grounding DINO: Marrying DINO with Grounded Pre- Training for Open-Set Object Detection, arXiv:2303.05499 [cs] (Mar. 2023). URL http://arxiv.org/abs/2303.05499

  58. [66]

    C. Xie, Z. Zhang, Y. Wu, F. Zhu, R. Zhao, S. Liang, Described Ob- ject Detection: Liberating Object Detection with Flexible Expressions, 31 arXiv:2307.12813 [cs] (Oct. 2023). URL http://arxiv.org/abs/2307.12813

  59. [67]

    Everingham, L

    M. Everingham, L. Van Gool, C. K. Williams, J. Winn, A. Zisserman, The pascal visual object classes (VOC) challenge, International journal of computer vision 88 (2009) 303–308

  60. [68]

    Pont-Tuset, L

    J. Pont-Tuset, L. Van Gool, Boosting object proposals: From PASCAL to COCO, in: Proceedings of the IEEE international conference on computer vision, 2015, pp. 1546–1554

  61. [69]

    B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, A. Torralba, Scene Parsing through ADE20K Dataset, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 633–641

  62. [70]

    Y. Liu, P. Sun, N. Wergeles, Y. Shang, A survey and performance evaluation of deep learning methods for small object detection, Expert Systems with Applications 172 (2021) 114602

  63. [71]

    van Erp, W

    M. van Erp, W. Tullett, V. Christlein, T. Ehrhart, A. H¨ urriyeto˘ glu, I. Lee- mans, P. Lisena, S. Menini, D. Schwabe, S. Tonelli, et al., More than the Name of the Rose: How to Make Computers Read, See, and Organize Smells, The American Historical Review 128 (1) (2023) 335–369

  64. [72]

    S. G. Magn´ usson, I. M. Szij´ art´ o, What is Microhistory?: Theory and Practice, Routledge, 2013

  65. [73]

    H. Ali, T. Paccosi, S. Menini, Z. Mathias, L. Pasquale, A. Kiymet, T. Rapha¨ el, M. van Erp, MUSTI-Multimodal Understanding of Smells in Texts and Images at MediaEval 2022, in: Proceedings of MediaEval 2022 CEUR Workshop, 2022

  66. [74]

    Howes, Sensual Relations: Engaging the Senses in Culture and Social Theory, University of Michiga/n Press, 2010

    D. Howes, Sensual Relations: Engaging the Senses in Culture and Social Theory, University of Michiga/n Press, 2010

  67. [75]

    Tullett, et al., Smell, History, and Heritage, The American Historical Review 127 (1) (2022) 261–309

    W. Tullett, et al., Smell, History, and Heritage, The American Historical Review 127 (1) (2022) 261–309

  68. [76]

    K. He, X. Zhang, S. Ren, J. Sun, Deep Residual Learning for Image Recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778

  69. [77]

    S. Xie, R. Girshick, P. Doll´ ar, Z. Tu, K. He, Aggregated Residual Transfor- mations for Deep Neural Networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1492–1500

  70. [78]

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, B. Guo, Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012–10022. 32

  71. [79]

    Ridnik, E

    T. Ridnik, E. Ben-Baruch, A. Noy, L. Zelnik-Manor, Imagenet-21k Pretrain- ing for the Masses, arXiv preprint arXiv:2104.10972 (2021)

  72. [80]

    J. Yang, C. Li, X. Dai, J. Gao, Focal Modulation Networks, Advances in Neural Information Processing Systems 35 (2022) 4203–4217

  73. [81]

    S. Ren, K. He, R. Girshick, J. Sun, Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Advances in neural information processing systems 28 (2015)

  74. [82]

    K. He, G. Gkioxari, P. Doll´ ar, R. Girshick, Mask R-CNN, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961–2969

  75. [83]

    T.-Y. Lin, P. Doll´ ar, R. Girshick, K. He, B. Hariharan, S. Belongie, Feature Pyramid Networks for Object Detection, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2117– 2125

  76. [84]

    Z. Cai, N. Vasconcelos, Cascade R-CNN: Delving into High Quality Object Detection, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6154–6162

  77. [85]

    Papers with Code, COCO object detection leaderboard, https://papers withcode.com/sota/object-detection-on-coco , accessed: February 19, 2023 (2021)

  78. [86]

    Zhang, F

    H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. Ni, H. Shum, DINO: DETR with Improved Denoising Anchor Boxes for End-to-End Object Detection, in: International Conference on Learning Representations, 2022

  79. [87]

    W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li, et al., InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions, arXiv preprint arXiv:2211.05778 (2022)

  80. [88]

    Carion, F

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, S. Zagoruyko, End-to-End Object Detection with Transformers, in: Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16, Springer, 2020, pp. 213–229

  81. [89]

    Z. Tian, C. Shen, H. Chen, T. He, FCOS: Fully Convolutional One-Stage Object Detection, in: 2019 IEEE/CVF International Conference on Com- puter Vision (ICCV), IEEE, Seoul, Korea (South), 2019, pp. 9626–9635. doi:10.1109/ICCV.2019.00972. URL https://ieeexplore.ieee.org/documen...

  82. [90]

    X. Xie, C. Lang, S. Miao, G. Cheng, K. Li, J. Han, Mutual-Assistance Learning for Object Detection, IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (12) (2023) 15171–15184. doi:10.1109/TPAMI.20 23.3319634. URL https://ieeexplore.ieee.org/document/10265160/ 33

  83. [91]

    Jocher, A

    G. Jocher, A. Chaurasia, J. Qiu, YOLO by Ultralytics, https://github.c om/ultralytics/ultralytics (1 2023). 34 Appendix A. Image Credits Image Credits Figure 1a Detail from: Still Life with Bouquet of Flowers . Jan Brueghel the Elder. 1610 –

  84. [1625]

    St¨ adel Museum, Frankfurt am Main

    Oil on copper. St¨ adel Museum, Frankfurt am Main. https://www.staedelm useum.de/go/ds/540. Figure 1b Detail from: Village in the Snow . David Teniers (II). 1625 – 1690. Oil on panel. RKD – Netherlands Institute for Art History, RKDImages(290572). Figure 1c Detail from: Man Sm...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.