Pith. sign in

REVIEW 2 major objections 1 minor 37 references

A lightweight SDF head on foundation features learns boundary distance maps from few masks to segment texture-poor industrial parts.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-26 14:51 UTC pith:MDSFWLBP

load-bearing objection The SDF-on-foundation-features idea targets a real industrial gap but the abstract supplies no numbers, so the performance claims stay untested. the 2 major comments →

arxiv 2606.21594 v1 pith:MDSFWLBP submitted 2026-06-19 cs.CV cs.RO

Boundary-by-Mask: Few-Shot Instance Segmentation with Mask-Conditioned Boundary Learning for Texture-Poor Industrial Parts

classification cs.CV cs.RO
keywords few-shot instance segmentationboundary learningsigned distance functiontexture-poor imagesindustrial partsmask conditioningfoundation model features
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces Boundary-by-Mask, a few-shot instance segmentation method that trains on boundaries rather than interior appearance. It extracts features from a foundation model encoder and fits a shallow MLP head to output signed distance maps conditioned on the provided instance masks. These maps are reconstructed into segmentation masks, enabling separation of instances even when surfaces have weak textures and uniform colors. The mask itself defines the target, so swapping it switches between whole-object and sub-part segmentation without retraining the head. Only the lightweight head is trained, allowing fast adaptation on a small number of examples from industrial and food domains.

Core claim

Boundary-by-Mask supervises a pixel-wise shallow MLP to predict signed distance function values from foundation-model features, conditioned on a few instance masks. The resulting boundary-aware distance maps are reconstructed into segmentation masks that separate instances by explicit contour estimation. This produces reliable masks on low-texture and color-uniform surfaces, while mask replacement directly controls the instance definition such as whole object versus sub-part.

What carries the argument

The mask-conditioned SDF head, a pixel-wise shallow MLP that maps foundation model features to boundary distance maps for SDF-to-mask reconstruction.

Load-bearing premise

Foundation-model features plus a few mask examples suffice to train the SDF head to output accurate boundary distance maps whose reconstruction yields correct instance masks in texture-poor regimes.

What would settle it

If reconstructed masks from the SDF predictions show large overlap errors or missed boundaries on a new collection of uniform-colored industrial parts, the reliability of boundary-supervised separation would be refuted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Explicit contour estimation separates instances reliably on surfaces with little texture or color contrast.
  • Replacing the conditioning mask changes the segmentation target without retraining the core model.
  • Rapid training is possible because only the shallow MLP head is fit on the limited examples.
  • The method generalizes to industrial parts and food items that have ambiguous boundaries.
  • Focus on boundaries rather than appearance maintains robustness when visual features are scarce.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The boundary-distance representation might support uncertainty maps around contours for quality control applications.
  • The same conditioning mechanism could be tested for consistent segmentation across image sequences or slight viewpoint changes.
  • Extending the reconstruction step to enforce topological constraints might reduce fragmentation on complex shapes.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript introduces Boundary-by-Mask, a few-shot instance segmentation framework for texture-poor industrial parts. It extracts features using a foundation-model encoder from a few RGB images and their instance masks, trains a lightweight Signed Distance Function (SDF) head (pixel-wise shallow MLP) to predict boundary-aware distance maps, and reconstructs segmentation masks from these maps. The approach allows conditioning the instance definition by the mask, enabling segmentation of whole objects or sub-parts. Experiments on industrial parts and food items are said to demonstrate strong few-shot generalization and robustness in feature-poor conditions.

Significance. If the performance claims hold, the work could have practical significance for industrial computer vision tasks where standard models struggle with low-texture and color-uniform surfaces. By supervising boundaries rather than interior appearance and using mask-conditioning, it offers a way to achieve reliable instance separation with minimal data and control over targets. The lightweight head for rapid training is a positive aspect.

major comments (2)
  1. [Abstract] Abstract: the abstract asserts strong few-shot generalization and robustness but supplies no quantitative results, baselines, error analysis, or dataset details; the central performance claims cannot be evaluated from the provided information.
  2. [Method] Method section: the framework assumes foundation-model features contain sufficient local contrast for the SDF head to regress accurate boundary distance maps in texture-poor regimes, but no feature analysis, ablation, or visualization is presented to show these features remain informative when textures and contextual cues are weak; this assumption is load-bearing for the claimed robustness.
minor comments (1)
  1. [Abstract] The abstract is lengthy and could be condensed while retaining the key claims.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments on our manuscript. We address each major point below and outline planned revisions.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the abstract asserts strong few-shot generalization and robustness but supplies no quantitative results, baselines, error analysis, or dataset details; the central performance claims cannot be evaluated from the provided information.

    Authors: We agree that the abstract would benefit from quantitative support. In the revised manuscript we will expand the abstract to include key metrics (e.g., mIoU and boundary F-score on the industrial-parts test set), a brief statement of the baselines used, and dataset characteristics so that the performance claims can be evaluated directly from the abstract. revision: yes

  2. Referee: [Method] Method section: the framework assumes foundation-model features contain sufficient local contrast for the SDF head to regress accurate boundary distance maps in texture-poor regimes, but no feature analysis, ablation, or visualization is presented to show these features remain informative when textures and contextual cues are weak; this assumption is load-bearing for the claimed robustness.

    Authors: The referee correctly notes the absence of explicit feature-level analysis. While the end-to-end results on texture-poor industrial and food-item data provide indirect support, we will add (i) side-by-side feature-map visualizations for low-texture versus textured inputs and (ii) an ablation replacing the foundation encoder with a randomly initialized or ImageNet-only backbone to quantify the contribution of pre-trained features. revision: yes

Circularity Check

0 steps flagged

No circularity; relies on external foundation models and standard supervised SDF training

full rationale

The provided abstract and description contain no equations, derivations, or self-citations that reduce claimed performance to fitted quantities or inputs defined by the method itself. The framework extracts features from an external foundation-model encoder and trains a lightweight MLP SDF head on given instance masks to regress distance maps, followed by reconstruction; this is standard supervised learning on external inputs rather than a self-referential loop. The instance definition being conditioned by the mask is an explicit design choice, not a circular reduction. The derivation is therefore self-contained against external benchmarks.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Abstract supplies no explicit free parameters, mathematical axioms, or newly postulated entities; ledger remains empty due to lack of technical detail.

pith-pipeline@v0.9.1-grok · 5745 in / 1047 out tokens · 21823 ms · 2026-06-26T14:51:02.617845+00:00 · methodology

0 comments
read the original abstract

Recent advances in large pre-trained models have led to remarkable progress in instance segmentation on general images. However, industrial scenarios remain challenging. Instance definitions are often application-specific and inconsistent, and the domain gap from general imagery is substantial due to weak textures and limited contextual cues. Consequently, a direct application of existing models is unreliable. We propose Boundary-by-Mask, a few-shot instance segmentation framework that supervises boundaries instead of interior appearance. Given a few RGB images and corresponding instance masks, the method extracts rich visual features using a foundation-model encoder and trains a lightweight Signed Distance Function (SDF) head to predict boundary-aware distance maps. Segmentation masks are obtained through an SDF-to-mask reconstruction process. By explicitly estimating contours, the framework achieves reliable instance separation even on low-texture and color-uniform surfaces. The instance definition is conditioned by the instance mask. Replacing the mask specifies the segmentation target, such as the whole object or a sub-part. A pixel-wise shallow MLP head enables rapid training. Experiments on industrial parts and food items with ambiguous boundaries show strong few-shot generalization, robustness in feature-poor conditions, and precise control over mask-level targets.

Figures

Figures reproduced from arXiv: 2606.21594 by Koichi Hashimoto, Naoya Chiba, Yutaka Yoshinaga.

Figure 1
Figure 1. Figure 1: With a few RGB–mask references, our Boundary Aware Pipeline is [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the method. The inputs are a few RGB–mask reference pairs and a query RGB image; the output is the instance mask for the query. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: SDF convention and example. Left: schematic—object boundary is the zero level; SDF > 0 inside and SDF < 0 outside. Middle: SDF of the tube objects (normalized signed distance; the zero level aligns with tube boundaries). Right: the original RGB image from which the SDF was computed. Ground-truth SDF Generation: Given a binary mask Mr , we first generate an SDF Dr that takes zero on the object boundary, has… view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative results of Boundary-by-Mask. Top : query RGB, middle : predicted SDF, bottom : output masks overlaid on the query. SAM2 YOLOv11 Pretrained YOLOv11 Finetuned PerSAM PerSAM-F No Time to Train! Input Image GT BbM(Ours) AP(50:95) - 0.80 0.00 0.00 0.70 0.00 0.00 0.00 AP(50:95) - 0.83 0.01 0.00 0.93 0.34 0.00 0.04 AP(50:95) - 0.54 0.21 0.07 0.48 0.05 0.05 0.18 [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative comparison on three queries (top: tubes, middle: fried chicken, bottom: screw pile). Left to right: input image, ground truth (GT), Boundary-by-Mask (BbM; ours), and baselines—SAM2 (AMG), YOLOv11 (pretrained), YOLOv11 (finetuned), PerSAM (no fine-tuning), PerSAM-F, and No time to train!. Shown AP values are per-image scores, computed independently for each image. TABLE I DATASET COMPOSITION (TE… view at source ↗
Figure 6
Figure 6. Figure 6: Mask-conditioned flexibility. By changing only the training targets for the SDF Head, the predicted instances adapt accordingly. Left to right: input image, object-level segmentation, part-level segmentation. chicken introduces irregular shapes with self-occlusion as a challenging test case beyond rigid parts. The dataset further covers flat and pile layouts with varying instance densities. Table I summari… view at source ↗
Figure 7
Figure 7. Figure 7: Accuracy vs. reference count K (COCO AP50:95). On our industrial-parts dataset, BbM (pink) rises steeply from K=1 to K=5 and then nearly saturates; it outperforms YOLOv11 for K ≤ 5, while YOLOv11 slightly overtakes at K=10. C. Experimental Results We first summarize the main findings for (i)–(iii), then point to figures/tables as supporting evidence. Representative BbM outputs are in [PITH_FULL_IMAGE:figu… view at source ↗
Figure 8
Figure 8. Figure 8: Offline reference-phase time for BbM as a function of [PITH_FULL_IMAGE:figures/full_fig_p007_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

37 extracted references · 7 canonical work pages · 5 internal anchors

  1. [1]

    Mask R-CNN,

    K. He, G. Gkioxari, P. Dollar, and R. Girshick, “Mask R-CNN,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2980–2988

  2. [2]

    End-to-End Object Detection with Transformers,

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-End Object Detection with Transformers,” in Proceedings of the European Conference on Computer Vision (ECCV), 2020, pp. 213–229

  3. [3]

    Segment Anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollar, and R. Girshick, “Segment Anything,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 3992– 4003

  4. [4]

    FGN: Fully Guided Network for Few-Shot Instance Segmentation,

    Z. Fan, J.-G. Yu, Z. Liang, J. Ou, C. Gao, G.-S. Xia, and Y . Li, “FGN: Fully Guided Network for Few-Shot Instance Segmentation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 9169–9178

  5. [5]

    Personalize Segment Anything Model with One Shot,

    R. Zhang, Z. Jiang, Z. Guo, S. Yan, J. Pan, H. Dong, Y . Qiao, P. Gao, and H. Li, “Personalize Segment Anything Model with One Shot,” inProceedings of the International Conference on Learning Representations (ICLR), 2024, pp. 18 250–18 279

  6. [6]

    PartNet: A Large-Scale Benchmark for Fine-Grained and Hi- erarchical Part-Level 3D Object Understanding,

    K. Mo, S. Zhu, A. X. Chang, L. Yi, S. Tripathi, L. J. Guibas, and H. Su, “PartNet: A Large-Scale Benchmark for Fine-Grained and Hi- erarchical Part-Level 3D Object Understanding,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 909–918

  7. [7]

    PANet: Few- Shot Image Semantic Segmentation With Prototype Alignment,

    K. Wang, J. H. Liew, Y . Zou, D. Zhou, and J. Feng, “PANet: Few- Shot Image Semantic Segmentation With Prototype Alignment,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 9196–9205

  8. [8]

    Edge Boxes: Locating Object Proposals from Edges,

    L. Zitnick and P. Dollar, “Edge Boxes: Locating Object Proposals from Edges,” inProceedings of the European Conference on Computer Vision (ECCV), 2014, pp. 391–405

  9. [9]

    Boundary loss for highly unbalanced segmentation,

    H. Kervadec, J. Bouchtiba, C. Desrosiers, E. Granger, J. Dolz, and I. Ben Ayed, “Boundary loss for highly unbalanced segmentation,” in Proceedings of the International Conference on Medical Imaging with Deep Learning (MIDL), vol. 102, 2019, pp. 285–296

  10. [10]

    Deep Watershed Transform for Instance Segmentation,

    M. Bai and R. Urtasun, “Deep Watershed Transform for Instance Segmentation,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2858–2866

  11. [11]

    One-Shot Instance Segmentation

    C. Michaelis, I. Ustyuzhaninov, M. Bethge, and A. S. Ecker, “One- Shot Instance Segmentation,”arXiv preprint arXiv:1811.11507, 2018

  12. [12]

    Incremental Few-Shot Instance Segmentation,

    D. A. Ganea, B. Boom, and R. Poppe, “Incremental Few-Shot Instance Segmentation,” inProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2021, pp. 1185–1194

  13. [13]

    Dynamic Transformer for Few-shot Instance Segmentation,

    H. Wang, J. Liu, Y . Liu, S. Maji, J.-J. Sonke, and E. Gavves, “Dynamic Transformer for Few-shot Instance Segmentation,” inProceedings of the ACM International Conference on Multimedia, 2022, pp. 2969– 2977

  14. [14]

    SAM-IF: Leveraging SAM for Incremental Few- Shot Instance Segmentation,

    X. Zhou and W. He, “SAM-IF: Leveraging SAM for Incremental Few- Shot Instance Segmentation,”arXiv preprint arXiv:2412.11034, 2024

  15. [15]

    Matcher: Segment Anything with One Shot Using All-Purpose Feature Match- ing,

    Y . Liu, M. Zhu, H. Li, H. Chen, X. Wang, and C. Shen, “Matcher: Segment Anything with One Shot Using All-Purpose Feature Match- ing,” inProceedings of the International Conference on Learning Representations (ICLR), 2024

  16. [16]

    No time to train! training-free reference-based instance segmentation.arXiv preprint arXiv:2507.02798, 2025

    M. Espinosa, C. Yang, L. Ericsson, S. McDonagh, and E. J. Crowley, “No time to train! Training-Free Reference-Based Instance Segmen- tation,”arXiv preprint arXiv:2507.02798, 2025

  17. [17]

    Emerging Properties in Self-Supervised Vision Trans- formers,

    M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging Properties in Self-Supervised Vision Trans- formers,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 9630–9640

  18. [18]

    DINOv2: Learning Robust Visual Features without Supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Je- gou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski, “DINOv2: Learning Robust Visual Features withou...

  19. [19]

    Sim ´eoni, H

    O. Sim ´eoni, H. V . V o, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. E. Yi, M. Ramamonjisoa, F. Massa, D. HAZIZA, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sen- tana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Couprie, J. Mairal, H. Jegou, P. Labatut, and P. Bojanowski, “DINOv3,”Transactions on Machine Le...

  20. [20]

    Masked Autoencoders Are Scalable Vision Learners,

    K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked Autoencoders Are Scalable Vision Learners,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 15 979–15 988

  21. [21]

    Masked-Attention Mask Transformer for Universal Image Segmen- tation,

    B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-Attention Mask Transformer for Universal Image Segmen- tation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 1280–1289

  22. [22]

    Segment Everything Everywhere All at Once,

    X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, and Y . J. Lee, “Segment Everything Everywhere All at Once,” in Advances in Neural Information Processing Systems (NeurIPS), 2023, pp. 19 769–19 782

  23. [23]

    SAM 2: Segment Anything in Images and Videos

    N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R ¨adle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V . Alwala, N. Carion, C.-Y . Wu, R. Girshick, P. Doll ´ar, and C. Feichtenhofer, “SAM 2: Segment Anything in Images and Videos,”arXiv preprint arXiv:2408.00714, 2024

  24. [24]

    SAM 3: Segment Anything with Concepts

    N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huang, J. Lei, T. Ma, B. Guo, A. Kalla, M. Marks, J. Greer, M. Wang, P. Sun, R. R ¨adle, T. Afouras, E. Mavroudi, K. Xu, T.-H. Wu, Y . Zhou, L. Momeni, R. Hazra, S. Ding, S. Vaze, F. Porcher, F. Li, S. Li, A. Kamath, H. K. Cheng, P. Doll ´ar, N. Ravi, K....

  25. [25]

    Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

    T. Ren, S. Liu, A. Zeng, J. Lin, K. Li, H. Cao, J. Chen, X. Huang, Y . Chen, F. Yan, Z. Zeng, H. Zhang, F. Li, J. Yang, H. Li, Q. Jiang, and L. Zhang, “Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks,”arXiv preprint 2401.14159, 2024

  26. [26]

    Spi- der: A Unified Framework for Context-dependent Concept Segmen- tation,

    X. Zhao, Y . Pang, W. Ji, B. Sheng, J. Zuo, L. Zhang, and H. Lu, “Spi- der: A Unified Framework for Context-dependent Concept Segmen- tation,” inProceedings of the International Conference on Machine Learning (ICML), 2024, pp. 60 906–60 926

  27. [27]

    Towards Robust Part-aware Instance Segmentation for Industrial Bin Picking,

    Y . Feng, B. Yang, X. Li, C.-W. Fu, R. Cao, K. Chen, Q. Dou, M. Wei, Y .-H. Liu, and P.-A. Heng, “Towards Robust Part-aware Instance Segmentation for Industrial Bin Picking,” inProceedings of the IEEE International Conference on Robotics and Automation (ICRA), 2022, pp. 405–411

  28. [28]

    Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion,

    Y . Zhu, N. Chiba, and K. Hashimoto, “Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion,” inProceedings of the British Machine Vision Conference (BMVC), 2025

  29. [29]

    Level set based shape prior segmentation,

    T. Chan. and W. Zhu, “Level set based shape prior segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2005, pp. 1164–1170

  30. [30]

    Euclidean distance mapping,

    P.-E. Danielsson, “Euclidean distance mapping,”Computer Graphics and Image Processing, vol. 14, no. 3, pp. 227–248, 1980

  31. [31]

    Watersheds in digital spaces: an efficient algorithm based on immersion simulations,

    L. M. Vincent and P. Soille, “Watersheds in digital spaces: an efficient algorithm based on immersion simulations,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 13, no. 6, pp. 583– 598, 1991

  32. [32]

    YOLOv11: An Overview of the Key Architectural Enhancements

    R. Khanam and M. Hussain, “YOLOv11: An Overview of the Key Architectural Enhancements,”arXiv preprint arXiv:2410.17725, 2024

  33. [33]

    Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts,

    X. Chen, R. Mottaghi, X. Liu, S. Fidler, R. Urtasun, and A. Yuille, “Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp. 1979–1986

  34. [34]

    PACO: Parts and Attributes of Common Objects,

    V . Ramanathan, A. Kalia, V . Petrovic, Y . Wen, B. Zheng, B. Guo, R. Wang, A. Marquez, R. Kovvuri, A. Kadian, A. Mousavi, Y . Song, A. Dubey, and D. Mahajan, “PACO: Parts and Attributes of Common Objects,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 7141–7151

  35. [35]

    Symmetry Aware Evaluation of 3D Object Detection and Pose Estimation in Scenes of Many Parts in Bulk,

    R. Bregier, F. Devernay, L. Leyrit, and J. L. Crowley, “Symmetry Aware Evaluation of 3D Object Detection and Pose Estimation in Scenes of Many Parts in Bulk,” inProceedings of the IEEE Interna- tional Conference on Computer Vision Workshops (ICCVW), 2017, pp. 2209–2218

  36. [36]

    Microsoft COCO: Common Objects in Context,

    T.-Y . Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Doll ´ar, “Microsoft COCO: Common Objects in Context,” inProceedings of the European Conference on Computer Vision (ECCV), 2014, pp. 740–755

  37. [37]

    Decoupled Weight Decay Regulariza- tion,

    I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regulariza- tion,” inProceedings of the International Conference on Learning Representations (ICLR), 2019