Pith. sign in

REVIEW 3 major objections 4 minor 216 references

Few-Shot Learning in Video and 3D Object Detection: A Survey

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims to be the first survey of few-shot detection in video and 3D, organized around tube proposals and temporal matching for video and prototype, reweighting, and incremental-branch heads on point clouds for 3D.

desk verdict A genuinely useful survey of few-shot video and 3D detection, but the citation tables are so unreliable that the current version cannot serve as a map of the field until they are corrected. read the letter →

arxiv 2507.17079 v1 pith:KGXHSOUX submitted 2025-07-22 cs.CV

classification cs.CV
keywords few-shotlearningvideoobjectdetection3Ddatascarcityprototypematchingpointcloudmeta-learningtubeproposalnetwork
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is a survey, and its claim is a map: few-shot learning for video and 3D object detection is a coherent field with a recognizable toolkit, and no earlier survey covered these two modalities together. The authors organize the video side around tube proposals, object trajectories stitched across frames, matched against a handful of support examples, with TPN and Thaw as the representative architectures. The 3D side is organized around point-cloud backbones teamed with prototype heads, class-specific reweighting vectors, and incremental branches for novel classes, with Prototypical VoteNet, MetaDet3D, and generalized few-shot 3D detection as the anchors. The survey also catalogs eleven open challenges, from class imbalance to temporal reasoning, and maps which methods already address each one. If the map is accurate, a researcher entering either domain gets a structured starting point instead of a scattered literature.

What carries the argument

The load-bearing objects are the spatiotemporal tube proposal and the few-shot matching head. A tube proposal is an object trajectory stitched across video frames, generated by a network like TPN and classified by a matching network after a temporal alignment step; it converts the annotation burden from labeling every frame to labeling a few support frames. On the 3D side the load-bearing objects are the prototype and reweighting heads placed on point-cloud backbones: geometric prototypes (class-agnostic, momentum-updated) and class prototypes (averaged support features) refined through cross-attention in Prototypical VoteNet; meta-learned class-specific reweighting vectors in MetaDet3D; and incremental classifier branches with a sample-adaptive balance loss in the generalized framework. These heads are what let a model compare sparse query points or tube features against a handful of support examples and transfer knowledge from base classes.

What would settle it

Open reference [30], which Table 1 credits as the source of the Tube Proposal Network: it is a Siamese-network paper on one-shot image classification and contains no tube proposal method, so a direct lookup settles that the table's attribution is wrong; the same check applied to Part-A2 Net's row in Table 2, which cites a video super-resolution paper, would confirm whether the survey's map can be trusted as a route to the primary literature.

Watch

Extended reading notes

Core claim

On the authors' own terms, the discovery is that few-shot detection generalizes to video and 3D data through a small number of recurring mechanisms. For video, the load-bearing idea is to replace single-frame detection with tube-level detection: a Tube Proposal Network generates spatiotemporal proposals that follow objects across frames, a Temporal Alignment Branch synchronizes query features, and a Tube Matching Network classifies the aggregated tube features against few-shot support images, so that temporal consistency substitutes for labeled data. For 3D, the load-bearing idea is to couple standard point-cloud backbones with few-shot heads: class-agnostic geometric prototypes and class-specific prototypes refine point and object features (Prototypical VoteNet), a meta-learned module produces class-specific reweighting vectors that guide voting and proposal generation (MetaDet3D), and frozen base networks with incremental novel-class branches plus an adaptive balance loss handle the base/novel imbalance (generalized FS3DOD). The survey's framing claim is that these mechanism-level recipes, taken together, are what make few-shot learning practical for surveillance and autonomous-driving deployment, and that the remaining barriers are the eleven open challenges it lists.

Load-bearing premise

The survey's value as a map depends on every method being credited to the paper that actually proposed it, and on every in-text citation agreeing with the tables; a reader who trusts the tables today could credit several methods to the wrong papers.

Editorial extensions

If this is right

  • If the survey's map is right, the strongest known recipe for few-shot video detection is two-stage: generate tube proposals, then classify them by matching aggregated tube features against support examples, with freeze-or-gradually-unfreeze fine-tuning and balanced sampling to avoid overfitting.
  • For few-shot 3D detection the corresponding recipe is a point-cloud backbone with a prototype, reweighting, or incremental-branch head, trained episodically with an adaptive loss that reweights positive, negative, and hard-negative samples.
  • The challenge-to-algorithm table gives newcomers a concrete agenda: methods already touch challenges like class imbalance and temporal reasoning, but no surveyed method addresses domain shift, scalability, interpretability, benchmarks, or hybrid combinations.
  • The survey's gap claim, if correct, means the two modalities were previously under-mapped, so this becomes the reference point for future surveys and for positioning new methods.
  • The information-theoretic caution in the conclusion implies that augmentation-based few-shot methods have a built-in ceiling: beyond the true information content of a few examples, synthesized data stops helping and starts misleading.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper mixes few-shot 3D detection proper with few-shot 3D action recognition (NGM and JEANIE); the editorial reading is that the two are converging on identical prototype-and-alignment machinery, a unification the survey itself never states.
  • Because several table rows cite work from neighboring tasks rather than few-shot detection proper, a reader should treat the tables as a scaffold: the field's core literature is smaller than the row counts, and the tables' real value is the mechanism-level taxonomy, which survives even where individual attributions need correction.
  • A testable extension follows from the survey's own information-theoretic caution: measure few-shot detection gains as a function of shots with and without advanced augmentation, and the returns should visibly plateau at one or two shots.
  • Because the video half rests on two representative architectures while the 3D half rests on several, the survey's own structure implies video few-shot detection is the younger subfield, which points toward a unified benchmark that evaluates the same novel classes in both modalities.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This survey reviews few-shot learning methods for video object detection and 3D object detection, motivated by the high cost of dense annotation in those modalities. It covers foundations of FSL and object detection, presents representative architectures such as TPN, Thaw, Prototypical VoteNet, generalized few-shot 3D detection, MetaDet3D, NGM Networks, and JEANIE, and closes with a list of open challenges and future directions. The paper's central claim is organizational: it aims to provide a comprehensive and accurate map of FSL methods for these two modalities, supported by comparison tables and cross-references to the literature.

Significance. If the map were reliable, the survey would fill a genuine gap: existing FSL surveys do not focus on video and 3D detection, and the paper collects several representative methods in one place. The qualitative descriptions of TPN, Thaw, Prototypical VoteNet, and MetaDet3D broadly match known published works, and the challenge taxonomy in Section 6 covers relevant issues such as class imbalance, temporal reasoning, and cross-domain transfer. However, the paper's utility as a citation authority is currently undermined by systematic misattributions in the tables and inconsistent in-text references, which are load-bearing for a survey of this kind.

major comments (3)
  1. [Tables 1 and 2; Sections 4.2.1, 4.3.3, 5.2.3] The tables contain systematic citation misattributions that break the survey's role as an accurate map. Table 1 pairs TPN with reference [30], which is the Siamese-networks classification paper of Koch et al., and Thaw with [178], which is the SPG 3D domain-adaptation paper. Table 2 pairs Part-A2 Net with [135], a video super-resolution paper, STEM-Seg with [66] (YOLOv4), and FSOD with [156] (MetaDet3D). A reader following these pointers would credit several methods to the wrong papers, so the central organizational claim is not supported in the current text.
  2. [Sections 4.2.1, 4.3.3, 5.2.5, 5.3, 6.12] In-text attributions are internally inconsistent. TPN is attributed to Fan et al. [119] in Section 4.2.1 but to [24] in Section 4.3.3; JEANIE is cited as [158] in Section 5.2.5 and as [159] in Section 5.3; and Section 6.12 refers to the '11 core challenges identified in Section 7,' although those challenges are actually presented in Section 6 and Section 7 contains the conclusions. The citation and cross-reference system therefore cannot be relied upon without independent verification.
  3. [Tables 1 and 3] The comparison tables mix few-shot video and 3D detection methods with methods from adjacent tasks (image few-shot detection, video super-resolution, semantic segmentation) without clearly distinguishing the out-of-scope rows. The caption of Table 1 even acknowledges 'including generic video object detection techniques,' but a reader of the tables cannot tell which entries are actual few-shot video detection methods and which are generic or from other tasks. Combined with the misattributions above, this weakens the survey's stated goal of providing a comprehensive and accurate comparison.
minor comments (4)
  1. [Section 3] The sentence 'It integrates ntegrates a Region Proposal Network' contains a typo and should read 'It integrates a Region Proposal Network.'
  2. [Section 4.1] The text 'This allows the FSV to quickly learn and generalize' appears to be truncated; 'FSV' is not defined and the sentence is incomplete.
  3. [Sections 4.2.1, 4.2.2, 5.2.1-5.2.3 and Supplementary Figures] The figure references are offset from the supplementary numbering: Section 4.2.1 refers to Figure S1 for TPN, while the supplementary's Figure S2 is the TPN diagram, and similar offsets occur for Thaw and the 3D detection figures. The cross-references should be aligned with the actual supplementary labels.
  4. [Section 7] The information-theoretic caveat about data augmentation is a useful addition, but it appears abruptly at the end of the conclusions without a precise statement or citation in the main text; the related references [203]-[205] appear only in the supplementary material.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey aggregates externally published methods and contains no derivation that reduces to its own inputs.

full rationale

This paper is a survey, so its central claims are organizational rather than derivational. It does not fit parameters to data and then rename them as predictions, and it does not invoke a uniqueness theorem or load-bearing self-citation to force its conclusions. The methods it discusses, such as TPN from Fan et al. [119], Thaw from [24], and Prototypical VoteNet from [154], are presented as external prior work with independent provenance. The citation inconsistencies in Tables 1 and 2 (e.g., TPN listed with [30], Part-A2 Net with [135], STEM-Seg with [66], and the internal conflict between [119] and [24] for TPN) are attribution-accuracy defects that would undermine the survey's reliability as a map of the literature, but they are not circular reasoning: a wrong pointer does not make the survey's summary equivalent to its own input. No equation, benchmark number, or qualitative verdict in the paper is derived from the paper's own assumptions in a way that would satisfy the circularity patterns defined for this analysis. Accordingly, the appropriate finding is no significant circularity, with a score of 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This survey introduces no free parameters or invented entities; all methods are attributed to prior work, even when the attributions are wrong. The analysis rests on two domain assumptions: that the coverage gap it claims is real, and that its representations of the cited methods and their performance are accurate. The second assumption is demonstrably violated in the comparison tables, which is the main epistemic risk of the paper.

assumptions (3)
  • domain assumption Prior FSL surveys have not focused specifically on video or 3D object detection, so this survey fills a gap.
    Asserted in Section 1.1 without an exhaustive comparison of prior surveys. The paper cites several general FSL surveys [7, 8, 20, 21, 22] but does not demonstrate that none of them covers video or 3D detection.
  • domain assumption The cited methods perform as their source papers report, and the survey's attributions are accurate.
    Load-bearing for claims such as 'Freeze attains the highest few-shot detection performance to date' (Section 4.2.2) and for the comparison tables. The text itself violates the attribution part, since Table 2 maps Part-A2 Net to reference [135] and STEM-Seg to reference [66].
  • standard math Standard object detection background (Faster R-CNN, YOLO, SSD) is correctly summarized.
    Section 3 restates textbook detector designs. The summaries are broadly correct, though they contain typos and are not original content.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Few-Shot Learning in Video and 3D Object Detection: A Survey." pith.science (2026). https://pith.science/paper/KGXHSOUX

@misc{pith2026250717079,
  author       = {Pith},
  title        = {Pith review of: Few-Shot Learning in Video and 3D Object Detection: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KGXHSOUX}},
  note         = {Machine review of arXiv:2507.17079}
}
read the original abstract

Few-shot learning (FSL) enables object detection models to recognize novel classes given only a few annotated examples, thereby reducing expensive manual data labeling. This survey examines recent FSL advances for video and 3D object detection. For video, FSL is especially valuable since annotating objects across frames is more laborious than for static images. By propagating information across frames, techniques like tube proposals and temporal matching networks can detect new classes from a couple examples, efficiently leveraging spatiotemporal structure. FSL for 3D detection from LiDAR or depth data faces challenges like sparsity and lack of texture. Solutions integrate FSL with specialized point cloud networks and losses tailored for class imbalance. Few-shot 3D detection enables practical autonomous driving deployment by minimizing costly 3D annotation needs. Core issues in both domains include balancing generalization and overfitting, integrating prototype matching, and handling data modality properties. In summary, FSL shows promise for reducing annotation requirements and enabling real-world video, 3D, and other applications by efficiently leveraging information across feature, temporal, and data modalities. By comprehensively surveying recent advancements, this paper illuminates FSL's potential to minimize supervision needs and enable deployment across video, 3D, and other real-world applications.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

216 extracted references · 63 canonical work pages

  1. [30]

    G. Koch, R. Zemel, R. Salakhutdinov, et al., Siamese neural networks for one-shot image recognition, in: ICML deep learning workshop, V ol. 2, Lille, 2015

  2. [178]

    Q. Xu, Y . Zhou, W. Wang, C. R. Qi, D. Anguelov, Spg: Unsupervised domain adaptation for 3d object detection via semantic point generation, in: Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, 2021, pp. 15446–15456

  3. [135]

    M. Liu, S. Jin, C. Yao, C. Lin, Y . Zhao, Temporal consis- tency learning of inter-frames for video super-resolution, IEEE Transactions on Circuits and Systems for Video Technology 33 (4) (2022) 1507–1520

  4. [66]

    Bochkovskiy, C.-Y

    A. Bochkovskiy, C.-Y . Wang, H.-Y . M. Liao, Yolov4: Optimal speed and accuracy of object detection, arXiv preprint arXiv:2004.10934 (2020)

  5. [156]

    S. Yuan, X. Li, H. Huang, Y . Fang, Meta-det3d: Learn to learn few-shot 3d object detection, in: Proceedings of the Asian Conference on Computer Vision, 2022, pp. 1761–1776

  6. [119]

    Fan, C.-K

    Q. Fan, C.-K. Tang, Y .-W. Tai, Few-shot video object detection, in: European Conference on Computer Vision, Springer, 2022, pp. 76–98

  7. [24]

    Z. Yu, G. Wang, L. Chen, S. Raschka, J. Luo, When few- shot learning meets video object detection, in: 2022 26th International Conference on Pattern Recognition (ICPR), IEEE, 2022, pp. 2986–2992

  8. [158]

    L. Wang, P. Koniusz, Temporal-viewpoint transportation plan for skeletal few-shot action recognition, in: Pro- ceedings of the Asian Conference on Computer Vision, 2022, pp. 4176–4193

  9. [159]

    L. Wang, J. Liu, P. Koniusz, 3d skeleton-based few-shot action recognition with jeanie is not so na \" ive, arXiv preprint arXiv:2112.12668 (2021)

Show all 216 references
  1. [1]

    T.-Y . Lin, P. Goyal, R. Girshick, K. He, P. Dollár, Focal loss for dense object detection, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 2980–2988

  2. [2]

    Krizhevsky, I

    A. Krizhevsky, I. Sutskever, G. E. Hinton, Imagenet classification with deep convolutional neural networks, Advances in neural information processing systems 25 (2012)

  3. [3]

    Z. Tian, C. Shen, H. Chen, T. He, Fcos: Fully convolu- tional one-stage object detection, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9627–9636

  4. [4]

    Alayrac, J

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al., Flamingo: a visual language model for few-shot learning, Advances in Neural Information Processing Systems 35 (2022) 23716–23736

  5. [5]

    F. Sung, Y . Yang, L. Zhang, T. Xiang, P. H. Torr, T. M. Hospedales, Learning to compare: Relation network for few-shot learning, in: Proceedings of the IEEE confer- ence on computer vision and pattern recognition, 2018, pp. 1199–1208

  6. [6]

    S. Ravi, H. Larochelle, Optimization as a model for few- shot learning, in: International conference on learning representations, 2016

  7. [7]

    Y . Wang, Q. Yao, J. T. Kwok, L. M. Ni, Generalizing from a few examples: A survey on few-shot learning, ACM computing surveys (csur) 53 (3) (2020) 1–34

  8. [8]

    Y . Song, T. Wang, P. Cai, S. K. Mondal, J. P. Sahoo, A comprehensive survey of few-shot learning: Evolution, applications, challenges, and opportunities, ACM Com- puting Surveys (2023)

  9. [9]

    Y . Xian, C. H. Lampert, B. Schiele, Z. Akata, Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly, IEEE transactions on pattern analysis and machine intelligence 41 (9) (2018) 2251–2265

  10. [10]

    Nichol, J

    A. Nichol, J. Achiam, J. Schulman, On first-order meta- learning algorithms, arXiv preprint arXiv:1803.02999 (2018)

  11. [11]

    Z. Li, F. Zhou, F. Chen, H. Li, Meta-sgd: Learning to learn quickly for few-shot learning, arXiv preprint arXiv:1707.09835 (2017)

  12. [12]

    Q. Sun, Y . Liu, T.-S. Chua, B. Schiele, Meta-transfer learning for few-shot learning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 403–412

  13. [13]

    Z. Yu, L. Chen, Z. Cheng, J. Luo, Transmatch: A transfer-learning scheme for semi-supervised few-shot learning, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2020, pp. 12856–12864

  14. [14]

    Jiang, K

    W. Jiang, K. Huang, J. Geng, X. Deng, Multi-scale met- ric learning for few-shot learning, IEEE Transactions on Circuits and Systems for Video Technology 31 (3) (2020) 1091–1102

  15. [15]

    L. Qiao, Y . Zhao, Z. Li, X. Qiu, J. Wu, C. Zhang, Defrcn: Decoupled faster r-cnn for few-shot object detection, in: Proceedings of the IEEE /CVF International Conference on Computer Vision, 2021, pp. 8681–8690

  16. [16]

    B. Sun, B. Li, S. Cai, Y . Yuan, C. Zhang, Fsce: Few-shot object detection via contrastive proposal encoding, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 7352–7362

  17. [17]

    G. Han, J. Ma, S. Huang, L. Chen, S.-F. Chang, Few-shot object detection with fully cross-transformer, in: Pro- ceedings of the IEEE /CVF conference on computer vi- sion and pattern recognition, 2022, pp. 5321–5330

  18. [18]

    Anderson, Video object recognition and detection, ht tps://www.itransition.com/blog/video-objec t-recognition-detection, accessed: 2023-08-31

    M. Anderson, Video object recognition and detection, ht tps://www.itransition.com/blog/video-objec t-recognition-detection, accessed: 2023-08-31

  19. [19]

    Kolesnikova, Detecting objects in video: a comprehen- sive guide 2022, https://mindtitan.com/resource s/blog/detecting-objects-in-video/ , accessed: 2023-08-31

    I. Kolesnikova, Detecting objects in video: a comprehen- sive guide 2022, https://mindtitan.com/resource s/blog/detecting-objects-in-video/ , accessed: 2023-08-31

  20. [20]

    Antonelli, D

    S. Antonelli, D. Avola, L. Cinque, D. Crisostomi, G. L. Foresti, F. Galasso, M. R. Marini, A. Mecca, D. Pannone, Few-shot object detection: A survey, ACM Computing Surveys (CSUR) 54 (11s) (2022) 1–37

  21. [21]

    Jiaxu, C

    L. Jiaxu, C. Taiyue, G. Xinbo, Y . Yongtao, W. Ye, G. Feng, W. Yue, A comparative review of recent few-shot object detection algorithms, arXiv preprint arXiv:2111.00201 (2021). 20

  22. [22]

    X. Li, Z. Sun, J.-H. Xue, Z. Ma, A concise review of re- cent few-shot meta-learning methods, Neurocomputing 456 (2021) 463–468

  23. [23]

    d’Archimbaud, Video Object Detection: AI’s New Challenge, https://kili-technology.com/data -labeling/computer-vision/video-annotatio n/video-object-detection, accessed: 2023-08-31

    E. d’Archimbaud, Video Object Detection: AI’s New Challenge, https://kili-technology.com/data -labeling/computer-vision/video-annotatio n/video-object-detection, accessed: 2023-08-31

  24. [25]

    Y . Zhou, O. Tuzel, V oxelnet: End-to-end learning for point cloud based 3d object detection, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4490–4499

  25. [26]

    C. R. Qi, L. Yi, H. Su, L. J. Guibas, Pointnet++: Deep hi- erarchical feature learning on point sets in a metric space, Advances in neural information processing systems 30 (2017)

  26. [27]

    Jiang, X

    Z. Jiang, X. Chen, X. Huang, X. Du, D. Zhou, Z. Wang, Back razor: Memory-e fficient transfer learning by self- sparsified backpropagation, Advances in Neural Infor- mation Processing Systems 35 (2022) 29248–29261

  27. [28]

    Snell, K

    J. Snell, K. Swersky, R. Zemel, Prototypical networks for few-shot learning, Advances in neural information pro- cessing systems 30 (2017)

  28. [29]

    C. Finn, P. Abbeel, S. Levine, Model-agnostic meta- learning for fast adaptation of deep networks, in: Inter- national conference on machine learning, PMLR, 2017, pp. 1126–1135

  29. [31]

    S. J. Pan, Q. Yang, A survey on transfer learning, IEEE Transactions on knowledge and data engineering 22 (10) (2009) 1345–1359

  30. [32]

    Boudiaf, I

    M. Boudiaf, I. Ziko, J. Rony, J. Dolz, P. Piantanida, I. Ben Ayed, Information maximization for few-shot learning, Advances in Neural Information Processing Systems 33 (2020) 2445–2457

  31. [33]

    Zhang, C

    J. Zhang, C. Zhao, B. Ni, M. Xu, X. Yang, Variational few-shot learning, in: Proceedings of the IEEE /CVF In- ternational Conference on Computer Vision, 2019, pp. 1685–1694

  32. [34]

    L. Qiao, Y . Shi, J. Li, Y . Wang, T. Huang, Y . Tian, Trans- ductive episodic-wise adaptive metric for few-shot learn- ing, in: Proceedings of the IEEE/CVF international con- ference on computer vision, 2019, pp. 3603–3612

  33. [35]

    Hajimiri, M

    S. Hajimiri, M. Boudiaf, I. Ben Ayed, J. Dolz, A strong baseline for generalized few-shot semantic segmenta- tion, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 11269–11278

  34. [36]

    Gavves, T

    E. Gavves, T. Mensink, T. Tommasi, C. G. Snoek, T. Tuytelaars, Active transfer learning with zero-shot pri- ors: Reusing past datasets for future tasks, in: Proceed- ings of the IEEE International Conference on Computer Vision, 2015, pp. 2731–2739

  35. [37]

    Mosbach, T

    M. Mosbach, T. Pimentel, S. Ravfogel, D. Klakow, Y . Elazar, Few-shot fine-tuning vs. in-context learn- ing: A fair comparison and evaluation, arXiv preprint arXiv:2305.16938 (2023)

  36. [38]

    Eustratiadis, Ł

    P. Eustratiadis, Ł. Dudziak, D. Li, T. Hospedales, Neural fine-tuning search for few-shot learning, arXiv preprint arXiv:2306.09295 (2023)

  37. [39]

    P. Peng, J. Wang, How to fine-tune deep neural networks in few-shot learning?, arXiv preprint arXiv:2012.00204 (2020)

  38. [40]

    H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, C. A. Ra ffel, Few-shot parameter-e fficient fine-tuning is better and cheaper than in-context learning, Advances in Neural Information Processing Systems 35 (2022) 1950–1965

  39. [41]

    Z. Shen, Z. Liu, J. Qin, M. Savvides, K.-T. Cheng, Partial is better than all: revisiting fine-tuning strategy for few- shot learning, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 35, 2021, pp. 9594–9602

  40. [42]

    S. X. Hu, D. Li, J. Stühmer, M. Kim, T. M. Hospedales, Pushing the limits of simple pipelines for few-shot learn- ing: External data and fine-tuning make a di fference, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 9068–9077

  41. [43]

    Y . Guo, H. Shi, A. Kumar, K. Grauman, T. Rosing, R. Feris, Spottune: transfer learning through adaptive fine-tuning, in: Proceedings of the IEEE /CVF confer- ence on computer vision and pattern recognition, 2019, pp. 4805–4814

  42. [44]

    C. Li, S. Li, H. Wang, F. Gu, A. D. Ball, Attention-based deep meta-transfer learning for few-shot fine-grained fault diagnosis, Knowledge-Based Systems 264 (2023) 110345

  43. [45]

    X. Xu, Z. Wang, Z. Chi, H. Yang, W. Du, Complemen- tary features based prototype self-updating for few-shot learning, Expert Systems with Applications 214 (2023) 119067

  44. [46]

    Shafahi, P

    A. Shafahi, P. Saadatpanah, C. Zhu, A. Ghiasi, C. Studer, D. Jacobs, T. Goldstein, Adversarially robust transfer learning, arXiv preprint arXiv:1905.08232 (2019). 21

  45. [47]

    E. Real, C. Liang, D. So, Q. Le, Automl-zero: Evolving machine learning algorithms from scratch, in: Interna- tional conference on machine learning, PMLR, 2020, pp. 8007–8019

  46. [48]

    Abuduweili, X

    A. Abuduweili, X. Li, H. Shi, C.-Z. Xu, D. Dou, Adaptive consistency regularization for semi-supervised transfer learning, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2021, pp. 6923–6932

  47. [49]

    D. Yan, J. Huang, H. Sun, F. Ding, Few-shot object de- tection with weight imprinting, Cognitive Computation (2023) 1–11

  48. [50]

    Hospedales, A

    T. Hospedales, A. Antoniou, P. Micaelli, A. Storkey, Meta-learning in neural networks: A survey, IEEE trans- actions on pattern analysis and machine intelligence 44 (9) (2021) 5149–5169

  49. [51]

    Vinyals, C

    O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra, et al., Matching networks for one shot learning, Advances in neural information processing systems 29 (2016)

  50. [52]

    Antoniou, A

    A. Antoniou, A. Storkey, H. Edwards, Data augmen- tation generative adversarial networks, arXiv preprint arXiv:1711.04340 (2017)

  51. [53]

    Shorten, T

    C. Shorten, T. M. Khoshgoftaar, A survey on image data augmentation for deep learning, Journal of Big Data 6 (1) (2019) 69

  52. [54]

    Lemley, S

    J. Lemley, S. Bazrafkan, P. Corcoran, Smart augmenta- tion learning an optimal data augmentation strategy, Ieee Access 5 (2017) 5858–5869

  53. [55]

    Goodfellow, Y

    I. Goodfellow, Y . Bengio, A. Courville, Regularization for deep learning, Deep learning (2016) 216–261

  54. [56]

    Srivastava, G

    N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, R. Salakhutdinov, Dropout: a simple way to prevent neu- ral networks from overfitting, The journal of machine learning research 15 (1) (2014) 1929–1958

  55. [57]

    Kuka ˇcka, V

    J. Kuka ˇcka, V . Golkov, D. Cremers, Regulariza- tion for deep learning: A taxonomy, arXiv preprint arXiv:1710.10686 (2017)

  56. [58]

    Y . Li, P. Zhang, X. Xu, Y . Lai, F. Shen, L. Chen, P. Gao, Few-shot prototype alignment regularization network for document image layout segementation, Pattern Recogni- tion 115 (2021) 107882

  57. [59]

    Girshick, J

    R. Girshick, J. Donahue, T. Darrell, J. Malik, Rich fea- ture hierarchies for accurate object detection and seman- tic segmentation, in: Proceedings of the IEEE confer- ence on computer vision and pattern recognition, 2014, pp. 580–587

  58. [60]

    Cheng, X

    G. Cheng, X. Yuan, X. Yao, K. Yan, Q. Zeng, X. Xie, J. Han, Towards large-scale small object detection: Sur- vey and benchmarks, IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)

  59. [61]

    Girshick, Fast r-cnn, in: Proceedings of the IEEE international conference on computer vision, 2015, pp

    R. Girshick, Fast r-cnn, in: Proceedings of the IEEE international conference on computer vision, 2015, pp. 1440–1448

  60. [62]

    S. Ren, K. He, R. Girshick, J. Sun, Faster r-cnn: Towards real-time object detection with region proposal networks, Advances in neural information processing systems 28 (2015)

  61. [63]

    Redmon, S

    J. Redmon, S. Divvala, R. Girshick, A. Farhadi, You only look once: Unified, real-time object detection, in: Pro- ceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 779–788

  62. [64]

    Redmon, A

    J. Redmon, A. Farhadi, Yolo9000: better, faster, stronger, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7263–7271

  63. [65]

    Redmon, A

    J. Redmon, A. Farhadi, Yolov3: An incremental im- provement, arXiv preprint arXiv:1804.02767 (2018)

  64. [67]

    YOLOv5 by Ultralytics, https://github.com/ultra lytics/yolov5, accessed: 2023-08-31

  65. [68]

    C. Li, L. Li, H. Jiang, K. Weng, Y . Geng, L. Li, Z. Ke, Q. Li, M. Cheng, W. Nie, et al., Yolov6: A single-stage object detection framework for industrial applications, arXiv preprint arXiv:2209.02976 (2022)

  66. [69]

    C.-Y . Wang, A. Bochkovskiy, H.-Y . M. Liao, Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 7464–7475

  67. [70]

    YOLOv8 by Ultralytics, https://github.com/ultra lytics/ultralytics, accessed: 2023-08-31

  68. [71]

    W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.- Y . Fu, A. C. Berg, Ssd: Single shot multibox detector, in: Computer Vision–ECCV 2016: 14th European Con- ference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14, Springer, 2016, pp. 21–37

  69. [72]

    F. He, N. Gao, Q. Li, S. Du, X. Zhao, K. Huang, Tempo- ral context enhanced feature aggregation for video object detection, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 34, 2020, pp. 10941–10948. 22

  70. [73]

    S. Lin, F. Qin, H. Peng, R. A. Bly, K. S. Moe, B. Han- naford, Multi-frame feature aggregation for real-time instrument segmentation in endoscopic video, IEEE Robotics and Automation Letters 6 (4) (2021) 6773– 6780

  71. [74]

    Cores, V

    D. Cores, V . M. Brea, M. Mucientes, Spatiotemporal tubelet feature aggregation and object linking for small object detection in videos, Applied Intelligence 53 (1) (2023) 1205–1217

  72. [75]

    X. Zhu, Y . Wang, J. Dai, L. Yuan, Y . Wei, Flow-guided feature aggregation for video object detection, in: Pro- ceedings of the IEEE international conference on com- puter vision, 2017, pp. 408–417

  73. [76]

    G. Sun, Y . Liu, H. Ding, T. Probst, L. Van Gool, Coarse- to-fine feature mining for video semantic segmentation, in: Proceedings of the IEEE /CVF Conference on Com- puter Vision and Pattern Recognition, 2022, pp. 3126– 3137

  74. [77]

    C. Xu, J. Zhang, M. Wang, G. Tian, Y . Liu, Multilevel spatial-temporal feature aggregation for video object de- tection, IEEE Transactions on Circuits and Systems for Video Technology 32 (11) (2022) 7809–7820

  75. [78]

    Honari, J

    S. Honari, J. Yosinski, P. Vincent, C. Pal, Recombinator networks: Learning coarse-to-fine feature aggregation, in: Proceedings of the IEEE conference on computer vi- sion and pattern recognition, 2016, pp. 5743–5752

  76. [79]

    J. Guo, W. Liu, S. Xin, Z. Zhao, B. Zhang, A frame level feature aggregation method for video target detection, in: 2021 33rd Chinese Control and Decision Conference (CCDC), IEEE, 2021, pp. 1368–1373

  77. [80]

    Muralidhara, K

    S. Muralidhara, K. A. Hashmi, A. Pagani, M. Liwicki, D. Stricker, M. Z. Afzal, Attention-guided disentangled feature aggregation for video object detection, Sensors 22 (21) (2022) 8583

  78. [81]

    L. Han, P. Wang, Z. Yin, F. Wang, H. Li, Exploiting better feature aggregation for video object detection, in: Proceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 1469–1477

  79. [82]

    Y . Qian, L. Yu, W. Liu, G. Kang, A. G. Hauptmann, Adaptive feature aggregation for video object detection, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, 2020, pp. 143–147

  80. [83]

    Q. Zhou, X. Li, L. He, Y . Yang, G. Cheng, Y . Tong, L. Ma, D. Tao, Transvod: end-to-end video object de- tection with spatial-temporal transformers, IEEE Trans- actions on Pattern Analysis and Machine Intelligence (2022)

  81. [84]

    Carion, F

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kir- illov, S. Zagoruyko, End-to-end object detection with transformers, in: European conference on computer vi- sion, Springer, 2020, pp. 213–229

  82. [85]

    Y . Wang, Z. Xu, X. Wang, C. Shen, B. Cheng, H. Shen, H. Xia, End-to-end video instance segmentation with transformers, in: Proceedings of the IEEE /CVF confer- ence on computer vision and pattern recognition, 2021, pp. 8741–8750

  83. [86]

    L. He, Q. Zhou, X. Li, L. Niu, G. Cheng, X. Li, W. Liu, Y . Tong, L. Ma, L. Zhang, End-to-end video object de- tection with spatial-temporal transformers, in: Proceed- ings of the 29th ACM International Conference on Mul- timedia, 2021, pp. 1507–1516

  84. [87]

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, B. Guo, Swin transformer: Hierarchical vision trans- former using shifted windows, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012–10022

  85. [88]

    Cui, Feature aggregated queries for transformer-based video object detectors, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2023, pp

    Y . Cui, Feature aggregated queries for transformer-based video object detectors, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2023, pp. 6365–6376

  86. [89]

    Cui, Faq: Feature aggregated queries for transformer-based video object detectors, arXiv preprint arXiv:2303.08319 (2023)

    Y . Cui, Faq: Feature aggregated queries for transformer-based video object detectors, arXiv preprint arXiv:2303.08319 (2023)

  87. [90]

    Z. Gao, Q. Wang, Z. Pan, Z. Zhai, H. Long, Pointpaint- ing: 3d object detection aided by semantic image infor- mation, Sensors 23 (5) (2023) 2868

  88. [91]

    S. Shi, X. Wang, H. Li, Pointrcnn: 3d object proposal generation and detection from point cloud, in: Proceed- ings of the IEEE /CVF conference on computer vision and pattern recognition, 2019, pp. 770–779

  89. [92]

    S. Shi, Z. Wang, X. Wang, H. Li, Part-a^ 2 net: 3d part- aware and aggregation neural network for object detec- tion from point cloud, arXiv preprint arXiv:1907.03670 2 (3) (2019)

  90. [93]

    Shreyas, M

    E. Shreyas, M. H. Sheth, et al., 3d object detection and tracking methods using deep learning for computer vi- sion applications, in: 2021 International Conference on Recent Trends on Electronics, Information, Communica- tion & Technology (RTEICT), IEEE, 2021, pp. 735–738

  91. [94]

    A. H. Lang, S. V ora, H. Caesar, L. Zhou, J. Yang, O. Bei- jbom, Pointpillars: Fast encoders for object detection from point clouds, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2019, pp. 12697–12705

  92. [95]

    Z. Tian, X. Chu, X. Wang, X. Wei, C. Shen, Fully convo- lutional one-stage 3d object detection on lidar range im- ages, Advances in Neural Information Processing Sys- tems 35 (2022) 34899–34911. 23

  93. [96]

    H. Zhao, M. Tian, S. Sun, J. Shao, J. Yan, S. Yi, X. Wang, X. Tang, Spindle net: Person re-identification with hu- man body region guided feature decomposition and fu- sion, in: Proceedings of the IEEE conference on com- puter vision and pattern recognition, 2017, pp. 1077– 1085

  94. [97]

    T. Yin, X. Zhou, P. Krahenbuhl, Center-based 3d object detection and tracking, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 11784–11793

  95. [98]

    H.-S. Kim, M. Won Lee, 3d object recognition using x3d and deep learning, in: The 25th International Conference on 3D Web Technology, 2020, pp. 1–8

  96. [99]

    X. Chen, K. Kundu, Z. Zhang, H. Ma, S. Fidler, R. Urta- sun, Monocular 3d object detection for autonomous driv- ing, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2147–2156

  97. [100]

    T. He, S. Soatto, Mono3d ++: Monocular 3d vehicle de- tection with two-scale 3d hypotheses and task priors, in: Proceedings of the AAAI Conference on Artificial Intel- ligence, V ol. 33, 2019, pp. 8409–8416

  98. [101]

    Y . You, Y . Wang, W.-L. Chao, D. Garg, G. Pleiss, B. Hariharan, M. Campbell, K. Q. Weinberger, Pseudo- lidar++: Accurate depth for 3d object detection in autonomous driving, arXiv preprint arXiv:1906.06310 (2019)

  99. [102]

    H. Liu, C. Wu, H. Wang, Real time object detection using lidar and camera fusion for autonomous driving, Scien- tific Reports 13 (1) (2023) 8056

  100. [103]

    J. Ku, M. Mozifian, J. Lee, A. Harakeh, S. L. Waslan- der, Joint 3d proposal generation and object detection from view aggregation, in: 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2018, pp. 1–8

  101. [104]

    V ora, A

    S. V ora, A. H. Lang, B. Helou, O. Beijbom, Pointpaint- ing: Sequential fusion for 3d object detection, in: Pro- ceedings of the IEEE /CVF conference on computer vi- sion and pattern recognition, 2020, pp. 4604–4612

  102. [105]

    D. Xu, D. Anguelov, A. Jain, Pointfusion: Deep sensor fusion for 3d bounding box estimation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 244–253

  103. [106]

    Belouadah, A

    E. Belouadah, A. Dapogny, K. Bailly, Multiod: Rehearsal-free multihead incremental object detector, arXiv preprint arXiv:2309.05334 (2023)

  104. [107]

    Krothapalli, L

    U. Krothapalli, L. Abbott, One size doesn’t fit all: Adap- tive label smoothing (2020)

  105. [108]

    W. Lv, S. Xu, Y . Zhao, G. Wang, J. Wei, C. Cui, Y . Du, Q. Dang, Y . Liu, Detrs beat yolos on real-time object detection, arXiv preprint arXiv:2304.08069 (2023)

  106. [109]

    H. Su, Y . He, R. Jiang, J. Zhang, W. Zou, B. Fan, Dsla: Dynamic smooth label assignment for e fficient anchor- free object detection, Pattern Recognition 131 (2022) 108868

  107. [110]

    T. Liu, L. Zhang, Y . Wang, J. Guan, Y . Fu, J. Zhao, S. Zhou, Recent few-shot object detection algorithms: A survey with performance comparison, ACM Trans- actions on Intelligent Systems and Technology 14 (4) (2023) 1–36

  108. [111]

    W. Jin, F. Guo, L. Zhu, Incremental self-supervised learning based on transformer for anomaly detection and localization, arXiv preprint arXiv:2303.17354 (2023)

  109. [112]

    Jiang, Z

    X. Jiang, Z. Li, M. Tian, J. Liu, S. Yi, D. Miao, Few-shot object detection via improved classification features, in: Proceedings of the IEEE /CVF Winter Conference on Applications of Computer Vision, 2023, pp. 5386–5395

  110. [113]

    Köhler, M

    M. Köhler, M. Eisenbach, H.-M. Gross, Few-shot object detection: a comprehensive survey, IEEE Transactions on Neural Networks and Learning Systems (2023)

  111. [114]

    A. Wu, S. Zhao, C. Deng, W. Liu, Generalized and dis- criminative few-shot object detection via svd-dictionary enhancement, Advances in Neural Information Process- ing Systems 34 (2021) 6353–6364

  112. [115]

    B. Kang, Z. Liu, X. Wang, F. Yu, J. Feng, T. Darrell, Few-shot object detection via feature reweighting, in: Proceedings of the IEEE /CVF International Conference on Computer Vision, 2019, pp. 8420–8429

  113. [116]

    Shangguan, M

    Z. Shangguan, M. Rostami, Identification of novel classes for improving few-shot object detection, arXiv preprint arXiv:2303.10422 (2023)

  114. [117]

    X. Wang, T. E. Huang, T. Darrell, J. E. Gonzalez, F. Yu, Frustratingly simple few-shot object detection, arXiv preprint arXiv:2003.06957 (2020)

  115. [118]

    J. Wang, D. Chen, Few-shot object detection method based on knowledge reasoning, Electronics 11 (9) (2022) 1327

  116. [120]

    Y . Chen, Y . Cao, H. Hu, L. Wang, Memory enhanced global-local aggregation for video object detection, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2020, pp. 10337–10346

  117. [121]

    K. Lee, H. Yang, S. Chakraborty, Z. Cai, G. Swami- nathan, A. Ravichandran, O. Dabeer, Rethinking few- shot object detection on a multi-domain benchmark, in: European Conference on Computer Vision, Springer, 2022, pp. 366–382. 24

  118. [122]

    Müller, S

    R. Müller, S. Kornblith, G. E. Hinton, When does label smoothing help?, Advances in neural information pro- cessing systems 32 (2019)

  119. [123]

    G. Han, Y . He, S. Huang, J. Ma, S.-F. Chang, Query adaptive few-shot object detection with heterogeneous graph convolutional networks, in: Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, 2021, pp. 3263–3272

  120. [124]

    Niklaus, L

    S. Niklaus, L. Mai, F. Liu, Video frame interpolation via adaptive separable convolution, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 261–270

  121. [125]

    M. Xu, S. Yoon, A. Fuentes, D. S. Park, A comprehen- sive survey of image augmentation techniques for deep learning, Pattern Recognition (2023) 109347

  122. [126]

    H. Wu, C. Song, S. Yue, Z. Wang, J. Xiao, Y . Liu, Dy- namic video mix-up for cross-domain action recognition, Neurocomputing 471 (2022) 358–368

  123. [127]

    Mangla, N

    P. Mangla, N. Kumari, A. Sinha, M. Singh, B. Krishna- murthy, V . N. Balasubramanian, Charting the right mani- fold: Manifold mixup for few-shot learning, in: Proceed- ings of the IEEE/CVF winter conference on applications of computer vision, 2020, pp. 2218–2227

  124. [128]

    A. Roy, A. Shah, K. Shah, P. Dhar, A. Cherian, R. Chel- lappa, Felmi: few shot learning with hard mixup, Ad- vances in Neural Information Processing Systems 35 (2022) 24474–24486

  125. [129]

    Nakamura, Y

    Y . Nakamura, Y . Ishii, Y . Maruyama, T. Yamashita, Few- shot adaptive object detection with cross-domain cutmix, in: Proceedings of the Asian Conference on Computer Vision, 2022, pp. 1350–1367

  126. [130]

    J. Yoo, N. Ahn, K.-A. Sohn, Rethinking data augmenta- tion for image super-resolution: A comprehensive analy- sis and a new strategy, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2020, pp. 8375–8384

  127. [131]

    Zhang, T

    C. Zhang, T. Yang, J. Weng, M. Cao, J. Wang, Y . Zou, Unsupervised pre-training for temporal action localiza- tion tasks, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 14031–14041

  128. [132]

    Aich, K.-C

    A. Aich, K.-C. Peng, A. K. Roy-Chowdhury, Cross- domain video anomaly detection without target domain adaptation, in: Proceedings of the IEEE /CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2579–2591

  129. [133]

    Olsson, W

    V . Olsson, W. Tranheden, J. Pinto, L. Svensson, Class- mix: Segmentation-based data augmentation for semi- supervised learning, in: Proceedings of the IEEE /CVF Winter Conference on Applications of Computer Vision, 2021, pp. 1369–1378

  130. [134]

    A. S. Chakravarthy, W.-D. Jang, Z. Lin, D. Wei, S. Bai, H. Pfister, Object propagation via inter-frame attentions for temporally stable video instance segmentation, arXiv preprint arXiv:2111.07529 (2021)

  131. [136]

    C. Deng, D. Chen, Q. Wu, Identity-consistent ag- gregation for video object detection, arXiv preprint arXiv:2308.07737 (2023)

  132. [137]

    Zhang, H

    Y . Zhang, H. Wang, H. Zhu, Z. Chen, Optical flow reusing for high-e fficiency space-time video super res- olution, IEEE Transactions on Circuits and Systems for Video Technology (2022)

  133. [138]

    J. Lin, X. Hu, Y . Cai, H. Wang, Y . Yan, X. Zou, Y . Zhang, L. Van Gool, Unsupervised flow-aligned sequence-to- sequence learning for video restoration, in: International Conference on Machine Learning, PMLR, 2022, pp. 13394–13404

  134. [139]

    X. Du, Y . Li, Y . Cui, R. Qian, J. Li, I. Bello, Revis- iting 3d resnets for video recognition, arXiv preprint arXiv:2109.01696 (2021)

  135. [140]

    Z. Ma, H. Zhang, J. Liu, Ms-lstm: Exploring spatiotem- poral multiscale representations in video prediction do- main, arXiv preprint arXiv:2304.07724 (2023)

  136. [141]

    P. L. Jeune, A. Mokraoui, A unified framework for attention-based few-shot object detection, arXiv preprint arXiv:2201.02052 (2022)

  137. [142]

    P. Pal, P. Chattopadhyay, M. Swarnkar, Temporal fea- ture aggregation with attention for insider threat detec- tion from activity logs, Expert Systems with Applica- tions 224 (2023) 119925

  138. [143]

    Gordevi ˇcius, J

    J. Gordevi ˇcius, J. Gamper, M. Böhlen, Parsimonious temporal aggregation, in: Proceedings of the 12th In- ternational Conference on Extending Database Technol- ogy: Advances in Database Technology, 2009, pp. 1006– 1017

  139. [144]

    Y . Fu, S. Sen, J. Reimann, C. Theurer, Spatiotempo- ral representation learning with gan trained lstm-lstm networks, in: 2020 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 2020, pp. 10548–10555

  140. [145]

    S. Dai, Y . Yu, H. Fan, J. Dong, Spatio-temporal represen- tation learning with social tie for personalized poi recom- mendation, Data Science and Engineering 7 (1) (2022) 44–56

  141. [146]

    Jin, Y .-F

    M. Jin, Y .-F. Li, Y . Zheng, B. Yang, S. Pan, Spatiotempo- ral representation learning on time series with dynamic graph odes (2021). 25

  142. [147]

    F. He, Q. Li, X. Zhao, K. Huang, Temporal-adaptive sparse feature aggregation for video object detection, Pattern Recognition 127 (2022) 108587

  143. [148]

    Nirthika, S

    R. Nirthika, S. Manivannan, A. Ramanan, R. Wang, Pooling in convolutional neural networks for medical image analysis: a survey and an empirical study, Neural Computing and Applications 34 (7) (2022) 5321–5347

  144. [149]

    X. Luo, X. Tu, Y . Ding, G. Gao, M. Deng, Expec- tation pooling: an e ffective and interpretable pooling method for predicting dna–protein binding, Bioinformat- ics 36 (5) (2020) 1405–1412

  145. [150]

    Rouvier, P.-M

    M. Rouvier, P.-M. Bousquet, J. Duret, Study on the tem- poral pooling used in deep neural networks for speaker verification, in: 2021 29th European Signal Processing Conference (EUSIPCO), IEEE, 2021, pp. 501–505

  146. [151]

    Zafar, M

    A. Zafar, M. Aamir, N. Mohd Nawi, A. Arshad, S. Riaz, A. Alruban, A. K. Dutta, S. Almotairi, A comparison of pooling methods for convolutional neural networks, Applied Sciences 12 (17) (2022) 8643

  147. [152]

    A.-K. N. Vu, N.-D. Nguyen, K.-D. Nguyen, V .-T. Nguyen, T. D. Ngo, T.-T. Do, T. V . Nguyen, Few-shot object detection via baby learning, Image and Vision Computing 120 (2022) 104398

  148. [153]

    Kim, H.-G

    G. Kim, H.-G. Jung, S.-W. Lee, Spatial reasoning for few-shot object detection, Pattern Recognition 120 (2021) 108118

  149. [154]

    S. Zhao, X. Qi, Prototypical votenet for few-shot 3d point cloud object detection, Advances in Neural Infor- mation Processing Systems 35 (2022) 13838–13851

  150. [155]

    J. Liu, X. Dong, S. Zhao, J. Shen, Generalized few-shot 3d object detection of lidar point cloud for autonomous driving, arXiv preprint arXiv:2302.03914 (2023)

  151. [157]

    M. Guo, E. Chou, D.-A. Huang, S. Song, S. Yeung, L. Fei-Fei, Neural graph matching networks for fewshot 3d action recognition, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 653– 669

  152. [160]

    F. Liu, S. Yang, D. Chen, H. Huang, J. Zhou, Few-shot classification guided by generalization error bound, Pat- tern Recognition (2023) 109904

  153. [161]

    J. He, Y . Chen, N. Wang, Z. Zhang, 3d video object de- tection with learnable object-centric global optimization, in: Proceedings of the IEEE /CVF Conference on Com- puter Vision and Pattern Recognition, 2023, pp. 5106– 5115

  154. [162]

    Brazil, G

    G. Brazil, G. Pons-Moll, X. Liu, B. Schiele, Kinematic 3d object detection in monocular video, in: Computer Vision–ECCV 2020: 16th European Conference, Glas- gow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, Springer, 2020, pp. 135–152

  155. [163]

    J. Wu, L. Song, T. Wang, Q. Zhang, J. Yuan, Forest r- cnn: Large-vocabulary long-tailed object detection and instance segmentation, in: Proceedings of the 28th ACM international conference on multimedia, 2020, pp. 1570– 1578

  156. [164]

    J. Wu, S. Liu, D. Huang, Y . Wang, Multi-scale posi- tive sample refinement for few-shot object detection, in: Computer Vision–ECCV 2020: 16th European Confer- ence, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVI 16, Springer, 2020, pp. 456–472

  157. [165]

    D. Lee, J. Kim, Resolving class imbalance for lidar- based object detector by dynamic weight average and contextual ground truth sampling, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Com- puter Vision, 2023, pp. 682–691

  158. [166]

    Z. Yu, G. Wang, L. Chen, S. Raschka, J. Luo, Few-shot learning for video object detection in a transfer-learning scheme, arXiv preprint arXiv:2103.14724 (2021)

  159. [167]

    S. Baik, M. Choi, J. Choi, H. Kim, K. M. Lee, Meta- learning with adaptive hyperparameters, Advances in neural information processing systems 33 (2020) 20755– 20765

  160. [168]

    Rajendran, A

    J. Rajendran, A. Irpan, E. Jang, Meta-learning requires meta-augmentation, Advances in Neural Information Processing Systems 33 (2020) 5705–5715

  161. [169]

    C. Si, X. Nie, W. Wang, L. Wang, T. Tan, J. Feng, Ad- versarial self-supervised learning for semi-supervised 3d action recognition, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part VII 16, Springer, 2020, pp. 35–51

  162. [170]

    G. Yang, D. Sun, V . Jampani, D. Vlasic, F. Cole, C. Liu, D. Ramanan, Viser: Video-specific surface em- beddings for articulated 3d shape reconstruction, Ad- vances in Neural Information Processing Systems 34 (2021) 19326–19338. 26

  163. [171]

    Huang, I

    G. Huang, I. Laradji, D. Vazquez, S. Lacoste-Julien, P. Rodriguez, A survey of self-supervised and few-shot object detection, IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (4) (2022) 4071–4089

  164. [172]

    F. Bao, G. Wu, C. Li, J. Zhu, B. Zhang, Stability and generalization of bilevel programming in hyperparam- eter optimization, Advances in neural information pro- cessing systems 34 (2021) 4529–4541

  165. [173]

    X. Luo, H. Wu, J. Zhang, L. Gao, J. Xu, J. Song, A closer look at few-shot classification again, arXiv preprint arXiv:2301.12246 (2023)

  166. [174]

    Zimmer, M

    W. Zimmer, M. Grabler, A. Knoll, Real-time and robust 3d object detection within road-side lidars using domain adaptation, arXiv preprint arXiv:2204.00132 (2022)

  167. [175]

    Y . Wang, J. Yin, W. Li, P. Frossard, R. Yang, J. Shen, Ssda3d: Semi-supervised domain adaptation for 3d ob- ject detection from point cloud, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 37, 2023, pp. 2707–2715

  168. [176]

    J. Yang, S. Shi, Z. Wang, H. Li, X. Qi, St3d: Self- training for unsupervised domain adaptation on 3d object detection, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2021, pp. 10368–10378

  169. [177]

    Hegde, V

    D. Hegde, V . Kilic, V . Sindagi, A. B. Cooper, M. Foster, V . M. Patel, Source-free unsupervised domain adaptation for 3d object detection in adverse weather, in: 2023 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 2023, pp. 6973–6980

  170. [179]

    J. Han, Y . Ren, J. Ding, K. Yan, G.-S. Xia, Few-shot ob- ject detection via variational feature aggregation, arXiv preprint arXiv:2301.13411 (2023)

  171. [180]

    I. H. Sarker, Machine learning: Algorithms, real-world applications and research directions, SN computer sci- ence 2 (3) (2021) 160

  172. [181]

    S. S. A. Zaidi, M. S. Ansari, A. Aslam, N. Kanwal, M. Asghar, B. Lee, A survey of modern deep learning based object detection models, Digital Signal Processing 126 (2022) 103514

  173. [182]

    H. Wang, X. Zhang, Y . Hu, Y . Yang, X. Cao, X. Zhen, Few-shot semantic segmentation with democratic atten- tion networks, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIII 16, Springer, 2020, pp. 730–746

  174. [183]

    C. Guo, B. Fan, Q. Zhang, S. Xiang, C. Pan, Augfpn: Im- proving multi-scale feature learning for object detection, in: Proceedings of the IEEE /CVF conference on com- puter vision and pattern recognition, 2020, pp. 12595– 12604

  175. [184]

    R. Cao, K. Zhang, Y . Chen, X. Yang, C. Jin, Point cloud completion via multi-scale edge convolution and atten- tion, in: Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 6183–6192

  176. [185]

    Z. Li, P. Xu, X. Chang, L. Yang, Y . Zhang, L. Yao, X. Chen, When object detection meets knowledge distil- lation: A survey, IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)

  177. [186]

    X. He, K. Zhao, X. Chu, Automl: A survey of the state-of-the-art, Knowledge-Based Systems 212 (2021) 106622

  178. [187]

    C. Chen, J. Wang, J. Pan, C. Bian, Z. Zhang, Graphskt: graph-guided structured knowledge transfer for domain adaptive lesion detection, IEEE Transactions on Medical Imaging 42 (2) (2022) 507–518

  179. [188]

    K. Liu, S. Lyu, P. Shivakumara, Y . Lu, Few-shot ob- ject segmentation with a new feature aggregation mod- ule, Displays 78 (2023) 102459

  180. [189]

    L. Zhao, G. Liu, D. Guo, W. Li, X. Fang, Boosting few- shot visual recognition via saliency-guided complemen- tary attention, Neurocomputing 507 (2022) 412–427

  181. [190]

    Munjal, A

    B. Munjal, A. Flaborea, S. Amin, F. Tombari, F. Galasso, Query-guided networks for few-shot fine-grained clas- sification and person search, Pattern Recognition 133 (2023) 109049

  182. [191]

    Mahapatra, Interpretable saliency maps and self- supervised learning for generalized zero shot medical image classification, arXiv preprint arXiv:2204.01728 (2022)

    D. Mahapatra, Interpretable saliency maps and self- supervised learning for generalized zero shot medical image classification, arXiv preprint arXiv:2204.01728 (2022)

  183. [192]

    W. Wang, L. Duan, Y . Wang, J. Fan, Z. Gong, Z. Zhang, A survey of deep visual cross-domain few-shot learning, arXiv preprint arXiv:2303.09253 (2023)

  184. [193]

    J. Zhu, J. Liu, S. Yang, Q. Zhang, X. He, Open bench- marking for click-through rate prediction, in: Proceed- ings of the 30th ACM International Conference on In- formation & Knowledge Management, 2021, pp. 2759– 2769

  185. [194]

    S. Sun, Y . Lu, S. Yu, X. Li, Z. Li, Z. Cao, Z. Liu, D. Ye, J. Bao, Rethinking dense retrieval’s few-shot abil- ity, arXiv preprint arXiv:2304.05845 (2023)

  186. [195]

    X. Wang, L. Lian, S. X. Yu, Unsupervised selective la- beling for more e ffective semi-supervised learning, in: European Conference on Computer Vision, Springer, 2022, pp. 427–445. 27

  187. [196]

    McClurg, A

    C. McClurg, A. Ayub, H. Tyagi, S. M. Rajtmajer, A. R. Wagner, Active class selection for few-shot class- incremental learning, arXiv preprint arXiv:2307.02641 (2023)

  188. [197]

    Z. Chen, J. Ge, H. Zhan, S. Huang, D. Wang, Pareto self- supervised training for few-shot learning, in: Proceed- ings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 13663–13672

  189. [198]

    Perez-Rua, X

    J.-M. Perez-Rua, X. Zhu, T. M. Hospedales, T. Xiang, Incremental few-shot object detection, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 13846–13855

  190. [199]

    Mousavian, D

    A. Mousavian, D. Anguelov, J. Flynn, J. Kosecka, 3d bounding box estimation using deep learning and geom- etry, in: Proceedings of the IEEE conference on Com- puter Vision and Pattern Recognition, 2017, pp. 7074– 7082

  191. [200]

    Manhardt, W

    F. Manhardt, W. Kehl, A. Gaidon, Roi-10d: Monocular lifting of 2d detection to 6d pose and metric shape, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 2069–2078

  192. [201]

    C. R. Qi, W. Liu, C. Wu, H. Su, L. J. Guibas, Frus- tum pointnets for 3d object detection from rgb-d data, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 918–927

  193. [202]

    X. Bai, Z. Hu, X. Zhu, Q. Huang, Y . Chen, H. Fu, C.-L. Tai, Transfusion: Robust lidar-camera fusion for 3d ob- ject detection with transformers, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 1090–1099

  194. [203]

    Wang, et al., Data augmentation for meta-learning, arXiv preprint arXiv:2002.08973 (2020)

    Y . Wang, et al., Data augmentation for meta-learning, arXiv preprint arXiv:2002.08973 (2020)

  195. [204]

    Franceschi, et al., A theoretical analysis of the num- ber of samples needed to estimate information-theoretic quantities, arXiv preprint arXiv:1708.01974 (2017)

    J.-Y . Franceschi, et al., A theoretical analysis of the num- ber of samples needed to estimate information-theoretic quantities, arXiv preprint arXiv:1708.01974 (2017)

  196. [205]

    Few-Shot Learning in Video and 3D Object Detection: A Survey

    E. D. Cubuk, et al., Randaugment: Practical data augmentation with no separate search, arXiv preprint arXiv:1909.13719 (2019). 28 The supplementary materials provide additional details and visual overviews to support the survey paper “Few-Shot Learning in Video and 3D Object D...

  197. [206]

    Foundations of Few-Shot Learning : Discusses key concepts like episodic training, problem formulations, meta-learning algorithms, metric-based approaches, data augmentation, and regularization techniques for few-shot learning

  198. [207]

    Foundations of Object Detection : Provides an overview of object detection methods, including two-stage and one-stage detectors, as well as video and 3D object detection approaches

  199. [208]

    Few-Shot Video Object Detection: Presents example frameworks and architectures tailored for few-shot video object detec- tion, highlighting techniques like metric learning, temporal feature aggregation, and episodic training

  200. [209]

    Few-Shot 3D Object Detection : Covers specialized few-shot detection methods for 3D data such as LiDAR point clouds, using techniques like geometric prototypes, support set guidance, and incremental learning. The supplementary materials expand on the key concepts, architecture...

  201. [210]

    learning to learn

    Foundations of Few-Shot Learning Few-shot learning (FSL) has emerged as a critical area of study within the deep learning framework, addressing one of the most pressing challenges in machine learning: the need for vast amounts of labeled data. In many real-world scenarios, obt...

  202. [211]

    Regions of Interest

    Foundations of Object Detection This section begins with an overview of object detection, including common techniques and applications. It then discusses various techniques used in video and 3D object detection. 3.1. Object Detection Object detection (OD) is a cornerstone of c...

  203. [212]

    and YOLO[63], which apply convolutional filters across an image in a single shot to directly output object locations and classes. The tradeoff between accuracy and speed makes two-stage detectors preferable for applications where accuracy is critical, while one-stage detectors...

  204. [213]

    Another line of work aggregates points into compact representations like pillars which encode vertical point columns, before applying e fficient 2D con- volutions on pseudo images

    extend it for 3D detection by first generating proposals which are then refined using point features. Another line of work aggregates points into compact representations like pillars which encode vertical point columns, before applying e fficient 2D con- volutions on pseudo im...

  205. [214]

    It first generates initial boxes from LiDAR, then fuses image features in the second decoder layer using soft-attention

    method fuses LiDAR and images using a novel transformer architecture with soft-attention. It first generates initial boxes from LiDAR, then fuses image features in the second decoder layer using soft-attention. This provides robustness to misalignment 34 and degraded image qua...

  206. [215]

    Few-Shot Video Object Detection 37 Figure S4: Overview of the Prototypical V oteNet architecture for few-shot 3D object detection [154]. It contains two key components - the Prototypical V ote Module (PVM) which refines local features using geometric prototypes, and the Protot...

  207. [216]

    It consists of a 3D Meta-Detector module that generates class-specific reweighting vectors zn from the few-shot support points

    Few-Shot 3D Object Detection 39 Figure S7: Overview of the MetaDet3D framework for few-shot 3D object detection [156]. It consists of a 3D Meta-Detector module that generates class-specific reweighting vectors zn from the few-shot support points. These reweighting vectors guid...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.