REVIEW 3 major objections 4 minor 216 references
Few-Shot Learning in Video and 3D Object Detection: A Survey
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims to be the first survey of few-shot detection in video and 3D, organized around tube proposals and temporal matching for video and prototype, reweighting, and incremental-branch heads on point clouds for 3D.
desk verdict A genuinely useful survey of few-shot video and 3D detection, but the citation tables are so unreliable that the current version cannot serve as a map of the field until they are corrected. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing objects are the spatiotemporal tube proposal and the few-shot matching head. A tube proposal is an object trajectory stitched across video frames, generated by a network like TPN and classified by a matching network after a temporal alignment step; it converts the annotation burden from labeling every frame to labeling a few support frames. On the 3D side the load-bearing objects are the prototype and reweighting heads placed on point-cloud backbones: geometric prototypes (class-agnostic, momentum-updated) and class prototypes (averaged support features) refined through cross-attention in Prototypical VoteNet; meta-learned class-specific reweighting vectors in MetaDet3D; and incremental classifier branches with a sample-adaptive balance loss in the generalized framework. These heads are what let a model compare sparse query points or tube features against a handful of support examples and transfer knowledge from base classes.
What would settle it
Open reference [30], which Table 1 credits as the source of the Tube Proposal Network: it is a Siamese-network paper on one-shot image classification and contains no tube proposal method, so a direct lookup settles that the table's attribution is wrong; the same check applied to Part-A2 Net's row in Table 2, which cites a video super-resolution paper, would confirm whether the survey's map can be trusted as a route to the primary literature.
Extended reading notes
Core claim
On the authors' own terms, the discovery is that few-shot detection generalizes to video and 3D data through a small number of recurring mechanisms. For video, the load-bearing idea is to replace single-frame detection with tube-level detection: a Tube Proposal Network generates spatiotemporal proposals that follow objects across frames, a Temporal Alignment Branch synchronizes query features, and a Tube Matching Network classifies the aggregated tube features against few-shot support images, so that temporal consistency substitutes for labeled data. For 3D, the load-bearing idea is to couple standard point-cloud backbones with few-shot heads: class-agnostic geometric prototypes and class-specific prototypes refine point and object features (Prototypical VoteNet), a meta-learned module produces class-specific reweighting vectors that guide voting and proposal generation (MetaDet3D), and frozen base networks with incremental novel-class branches plus an adaptive balance loss handle the base/novel imbalance (generalized FS3DOD). The survey's framing claim is that these mechanism-level recipes, taken together, are what make few-shot learning practical for surveillance and autonomous-driving deployment, and that the remaining barriers are the eleven open challenges it lists.
Load-bearing premise
The survey's value as a map depends on every method being credited to the paper that actually proposed it, and on every in-text citation agreeing with the tables; a reader who trusts the tables today could credit several methods to the wrong papers.
Editorial extensions
If this is right
- If the survey's map is right, the strongest known recipe for few-shot video detection is two-stage: generate tube proposals, then classify them by matching aggregated tube features against support examples, with freeze-or-gradually-unfreeze fine-tuning and balanced sampling to avoid overfitting.
- For few-shot 3D detection the corresponding recipe is a point-cloud backbone with a prototype, reweighting, or incremental-branch head, trained episodically with an adaptive loss that reweights positive, negative, and hard-negative samples.
- The challenge-to-algorithm table gives newcomers a concrete agenda: methods already touch challenges like class imbalance and temporal reasoning, but no surveyed method addresses domain shift, scalability, interpretability, benchmarks, or hybrid combinations.
- The survey's gap claim, if correct, means the two modalities were previously under-mapped, so this becomes the reference point for future surveys and for positioning new methods.
- The information-theoretic caution in the conclusion implies that augmentation-based few-shot methods have a built-in ceiling: beyond the true information content of a few examples, synthesized data stops helping and starts misleading.
Reading between the lines
- The paper mixes few-shot 3D detection proper with few-shot 3D action recognition (NGM and JEANIE); the editorial reading is that the two are converging on identical prototype-and-alignment machinery, a unification the survey itself never states.
- Because several table rows cite work from neighboring tasks rather than few-shot detection proper, a reader should treat the tables as a scaffold: the field's core literature is smaller than the row counts, and the tables' real value is the mechanism-level taxonomy, which survives even where individual attributions need correction.
- A testable extension follows from the survey's own information-theoretic caution: measure few-shot detection gains as a function of shots with and without advanced augmentation, and the returns should visibly plateau at one or two shots.
- Because the video half rests on two representative architectures while the 3D half rests on several, the survey's own structure implies video few-shot detection is the younger subfield, which points toward a unified benchmark that evaluates the same novel classes in both modalities.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews few-shot learning methods for video object detection and 3D object detection, motivated by the high cost of dense annotation in those modalities. It covers foundations of FSL and object detection, presents representative architectures such as TPN, Thaw, Prototypical VoteNet, generalized few-shot 3D detection, MetaDet3D, NGM Networks, and JEANIE, and closes with a list of open challenges and future directions. The paper's central claim is organizational: it aims to provide a comprehensive and accurate map of FSL methods for these two modalities, supported by comparison tables and cross-references to the literature.
Significance. If the map were reliable, the survey would fill a genuine gap: existing FSL surveys do not focus on video and 3D detection, and the paper collects several representative methods in one place. The qualitative descriptions of TPN, Thaw, Prototypical VoteNet, and MetaDet3D broadly match known published works, and the challenge taxonomy in Section 6 covers relevant issues such as class imbalance, temporal reasoning, and cross-domain transfer. However, the paper's utility as a citation authority is currently undermined by systematic misattributions in the tables and inconsistent in-text references, which are load-bearing for a survey of this kind.
major comments (3)
- [Tables 1 and 2; Sections 4.2.1, 4.3.3, 5.2.3] The tables contain systematic citation misattributions that break the survey's role as an accurate map. Table 1 pairs TPN with reference [30], which is the Siamese-networks classification paper of Koch et al., and Thaw with [178], which is the SPG 3D domain-adaptation paper. Table 2 pairs Part-A2 Net with [135], a video super-resolution paper, STEM-Seg with [66] (YOLOv4), and FSOD with [156] (MetaDet3D). A reader following these pointers would credit several methods to the wrong papers, so the central organizational claim is not supported in the current text.
- [Sections 4.2.1, 4.3.3, 5.2.5, 5.3, 6.12] In-text attributions are internally inconsistent. TPN is attributed to Fan et al. [119] in Section 4.2.1 but to [24] in Section 4.3.3; JEANIE is cited as [158] in Section 5.2.5 and as [159] in Section 5.3; and Section 6.12 refers to the '11 core challenges identified in Section 7,' although those challenges are actually presented in Section 6 and Section 7 contains the conclusions. The citation and cross-reference system therefore cannot be relied upon without independent verification.
- [Tables 1 and 3] The comparison tables mix few-shot video and 3D detection methods with methods from adjacent tasks (image few-shot detection, video super-resolution, semantic segmentation) without clearly distinguishing the out-of-scope rows. The caption of Table 1 even acknowledges 'including generic video object detection techniques,' but a reader of the tables cannot tell which entries are actual few-shot video detection methods and which are generic or from other tasks. Combined with the misattributions above, this weakens the survey's stated goal of providing a comprehensive and accurate comparison.
minor comments (4)
- [Section 3] The sentence 'It integrates ntegrates a Region Proposal Network' contains a typo and should read 'It integrates a Region Proposal Network.'
- [Section 4.1] The text 'This allows the FSV to quickly learn and generalize' appears to be truncated; 'FSV' is not defined and the sentence is incomplete.
- [Sections 4.2.1, 4.2.2, 5.2.1-5.2.3 and Supplementary Figures] The figure references are offset from the supplementary numbering: Section 4.2.1 refers to Figure S1 for TPN, while the supplementary's Figure S2 is the TPN diagram, and similar offsets occur for Thaw and the 3D detection figures. The cross-references should be aligned with the actual supplementary labels.
- [Section 7] The information-theoretic caveat about data augmentation is a useful addition, but it appears abruptly at the end of the conclusions without a precise statement or citation in the main text; the related references [203]-[205] appear only in the supplementary material.
Circularity Check
No significant circularity: the survey aggregates externally published methods and contains no derivation that reduces to its own inputs.
full rationale
This paper is a survey, so its central claims are organizational rather than derivational. It does not fit parameters to data and then rename them as predictions, and it does not invoke a uniqueness theorem or load-bearing self-citation to force its conclusions. The methods it discusses, such as TPN from Fan et al. [119], Thaw from [24], and Prototypical VoteNet from [154], are presented as external prior work with independent provenance. The citation inconsistencies in Tables 1 and 2 (e.g., TPN listed with [30], Part-A2 Net with [135], STEM-Seg with [66], and the internal conflict between [119] and [24] for TPN) are attribution-accuracy defects that would undermine the survey's reliability as a map of the literature, but they are not circular reasoning: a wrong pointer does not make the survey's summary equivalent to its own input. No equation, benchmark number, or qualitative verdict in the paper is derived from the paper's own assumptions in a way that would satisfy the circularity patterns defined for this analysis. Accordingly, the appropriate finding is no significant circularity, with a score of 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Prior FSL surveys have not focused specifically on video or 3D object detection, so this survey fills a gap.
- domain assumption The cited methods perform as their source papers report, and the survey's attributions are accurate.
- standard math Standard object detection background (Faster R-CNN, YOLO, SSD) is correctly summarized.
Cite this review
Pith. "Pith review of Few-Shot Learning in Video and 3D Object Detection: A Survey." pith.science (2026). https://pith.science/paper/KGXHSOUX
@misc{pith2026250717079,
author = {Pith},
title = {Pith review of: Few-Shot Learning in Video and 3D Object Detection: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/KGXHSOUX}},
note = {Machine review of arXiv:2507.17079}
}
read the original abstract
Few-shot learning (FSL) enables object detection models to recognize novel classes given only a few annotated examples, thereby reducing expensive manual data labeling. This survey examines recent FSL advances for video and 3D object detection. For video, FSL is especially valuable since annotating objects across frames is more laborious than for static images. By propagating information across frames, techniques like tube proposals and temporal matching networks can detect new classes from a couple examples, efficiently leveraging spatiotemporal structure. FSL for 3D detection from LiDAR or depth data faces challenges like sparsity and lack of texture. Solutions integrate FSL with specialized point cloud networks and losses tailored for class imbalance. Few-shot 3D detection enables practical autonomous driving deployment by minimizing costly 3D annotation needs. Core issues in both domains include balancing generalization and overfitting, integrating prototype matching, and handling data modality properties. In summary, FSL shows promise for reducing annotation requirements and enabling real-world video, 3D, and other applications by efficiently leveraging information across feature, temporal, and data modalities. By comprehensively surveying recent advancements, this paper illuminates FSL's potential to minimize supervision needs and enable deployment across video, 3D, and other real-world applications.
Reference graph
Works this paper leans on
-
[30]
G. Koch, R. Zemel, R. Salakhutdinov, et al., Siamese neural networks for one-shot image recognition, in: ICML deep learning workshop, V ol. 2, Lille, 2015
2015
-
[178]
Q. Xu, Y . Zhou, W. Wang, C. R. Qi, D. Anguelov, Spg: Unsupervised domain adaptation for 3d object detection via semantic point generation, in: Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, 2021, pp. 15446–15456
work page 2021
-
[135]
M. Liu, S. Jin, C. Yao, C. Lin, Y . Zhao, Temporal consis- tency learning of inter-frames for video super-resolution, IEEE Transactions on Circuits and Systems for Video Technology 33 (4) (2022) 1507–1520
2022
-
[66]
A. Bochkovskiy, C.-Y . Wang, H.-Y . M. Liao, Yolov4: Optimal speed and accuracy of object detection, arXiv preprint arXiv:2004.10934 (2020)
arXiv 2020
-
[156]
S. Yuan, X. Li, H. Huang, Y . Fang, Meta-det3d: Learn to learn few-shot 3d object detection, in: Proceedings of the Asian Conference on Computer Vision, 2022, pp. 1761–1776
2022
-
[119]
Fan, C.-K
Q. Fan, C.-K. Tang, Y .-W. Tai, Few-shot video object detection, in: European Conference on Computer Vision, Springer, 2022, pp. 76–98
2022
-
[24]
Z. Yu, G. Wang, L. Chen, S. Raschka, J. Luo, When few- shot learning meets video object detection, in: 2022 26th International Conference on Pattern Recognition (ICPR), IEEE, 2022, pp. 2986–2992
2022
-
[158]
L. Wang, P. Koniusz, Temporal-viewpoint transportation plan for skeletal few-shot action recognition, in: Pro- ceedings of the Asian Conference on Computer Vision, 2022, pp. 4176–4193
2022
-
[159]
L. Wang, J. Liu, P. Koniusz, 3d skeleton-based few-shot action recognition with jeanie is not so na \" ive, arXiv preprint arXiv:2112.12668 (2021)
work page Pith review arXiv 2021
Show all 216 references
-
[1]
T.-Y . Lin, P. Goyal, R. Girshick, K. He, P. Dollár, Focal loss for dense object detection, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 2980–2988
2017
-
[2]
Krizhevsky, I
A. Krizhevsky, I. Sutskever, G. E. Hinton, Imagenet classification with deep convolutional neural networks, Advances in neural information processing systems 25 (2012)
2012
-
[3]
Z. Tian, C. Shen, H. Chen, T. He, Fcos: Fully convolu- tional one-stage object detection, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9627–9636
2019
-
[4]
Alayrac, J
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al., Flamingo: a visual language model for few-shot learning, Advances in Neural Information Processing Systems 35 (2022) 23716–23736
2022
-
[5]
F. Sung, Y . Yang, L. Zhang, T. Xiang, P. H. Torr, T. M. Hospedales, Learning to compare: Relation network for few-shot learning, in: Proceedings of the IEEE confer- ence on computer vision and pattern recognition, 2018, pp. 1199–1208
2018
-
[6]
S. Ravi, H. Larochelle, Optimization as a model for few- shot learning, in: International conference on learning representations, 2016
2016
-
[7]
Y . Wang, Q. Yao, J. T. Kwok, L. M. Ni, Generalizing from a few examples: A survey on few-shot learning, ACM computing surveys (csur) 53 (3) (2020) 1–34
2020
-
[8]
Y . Song, T. Wang, P. Cai, S. K. Mondal, J. P. Sahoo, A comprehensive survey of few-shot learning: Evolution, applications, challenges, and opportunities, ACM Com- puting Surveys (2023)
2023
-
[9]
Y . Xian, C. H. Lampert, B. Schiele, Z. Akata, Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly, IEEE transactions on pattern analysis and machine intelligence 41 (9) (2018) 2251–2265
2018
-
[10]
Nichol, J
A. Nichol, J. Achiam, J. Schulman, On first-order meta- learning algorithms, arXiv preprint arXiv:1803.02999 (2018)
2018 arXiv
-
[11]
Z. Li, F. Zhou, F. Chen, H. Li, Meta-sgd: Learning to learn quickly for few-shot learning, arXiv preprint arXiv:1707.09835 (2017)
2017 arXiv
-
[12]
Q. Sun, Y . Liu, T.-S. Chua, B. Schiele, Meta-transfer learning for few-shot learning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 403–412
2019
-
[13]
Z. Yu, L. Chen, Z. Cheng, J. Luo, Transmatch: A transfer-learning scheme for semi-supervised few-shot learning, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2020, pp. 12856–12864
2020
-
[14]
Jiang, K
W. Jiang, K. Huang, J. Geng, X. Deng, Multi-scale met- ric learning for few-shot learning, IEEE Transactions on Circuits and Systems for Video Technology 31 (3) (2020) 1091–1102
2020
-
[15]
L. Qiao, Y . Zhao, Z. Li, X. Qiu, J. Wu, C. Zhang, Defrcn: Decoupled faster r-cnn for few-shot object detection, in: Proceedings of the IEEE /CVF International Conference on Computer Vision, 2021, pp. 8681–8690
2021
-
[16]
B. Sun, B. Li, S. Cai, Y . Yuan, C. Zhang, Fsce: Few-shot object detection via contrastive proposal encoding, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 7352–7362
2021
-
[17]
G. Han, J. Ma, S. Huang, L. Chen, S.-F. Chang, Few-shot object detection with fully cross-transformer, in: Pro- ceedings of the IEEE /CVF conference on computer vi- sion and pattern recognition, 2022, pp. 5321–5330
2022
-
[18]
Anderson, Video object recognition and detection, ht tps://www.itransition.com/blog/video-objec t-recognition-detection, accessed: 2023-08-31
M. Anderson, Video object recognition and detection, ht tps://www.itransition.com/blog/video-objec t-recognition-detection, accessed: 2023-08-31
2023
-
[19]
Kolesnikova, Detecting objects in video: a comprehen- sive guide 2022, https://mindtitan.com/resource s/blog/detecting-objects-in-video/ , accessed: 2023-08-31
I. Kolesnikova, Detecting objects in video: a comprehen- sive guide 2022, https://mindtitan.com/resource s/blog/detecting-objects-in-video/ , accessed: 2023-08-31
2022
-
[20]
Antonelli, D
S. Antonelli, D. Avola, L. Cinque, D. Crisostomi, G. L. Foresti, F. Galasso, M. R. Marini, A. Mecca, D. Pannone, Few-shot object detection: A survey, ACM Computing Surveys (CSUR) 54 (11s) (2022) 1–37
2022
-
[21]
Jiaxu, C
L. Jiaxu, C. Taiyue, G. Xinbo, Y . Yongtao, W. Ye, G. Feng, W. Yue, A comparative review of recent few-shot object detection algorithms, arXiv preprint arXiv:2111.00201 (2021). 20
2021 arXiv
-
[22]
X. Li, Z. Sun, J.-H. Xue, Z. Ma, A concise review of re- cent few-shot meta-learning methods, Neurocomputing 456 (2021) 463–468
2021
-
[23]
d’Archimbaud, Video Object Detection: AI’s New Challenge, https://kili-technology.com/data -labeling/computer-vision/video-annotatio n/video-object-detection, accessed: 2023-08-31
E. d’Archimbaud, Video Object Detection: AI’s New Challenge, https://kili-technology.com/data -labeling/computer-vision/video-annotatio n/video-object-detection, accessed: 2023-08-31
2023
-
[25]
Y . Zhou, O. Tuzel, V oxelnet: End-to-end learning for point cloud based 3d object detection, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4490–4499
2018
-
[26]
C. R. Qi, L. Yi, H. Su, L. J. Guibas, Pointnet++: Deep hi- erarchical feature learning on point sets in a metric space, Advances in neural information processing systems 30 (2017)
2017
-
[27]
Jiang, X
Z. Jiang, X. Chen, X. Huang, X. Du, D. Zhou, Z. Wang, Back razor: Memory-e fficient transfer learning by self- sparsified backpropagation, Advances in Neural Infor- mation Processing Systems 35 (2022) 29248–29261
2022
-
[28]
Snell, K
J. Snell, K. Swersky, R. Zemel, Prototypical networks for few-shot learning, Advances in neural information pro- cessing systems 30 (2017)
2017
-
[29]
C. Finn, P. Abbeel, S. Levine, Model-agnostic meta- learning for fast adaptation of deep networks, in: Inter- national conference on machine learning, PMLR, 2017, pp. 1126–1135
2017
-
[31]
S. J. Pan, Q. Yang, A survey on transfer learning, IEEE Transactions on knowledge and data engineering 22 (10) (2009) 1345–1359
2009
-
[32]
Boudiaf, I
M. Boudiaf, I. Ziko, J. Rony, J. Dolz, P. Piantanida, I. Ben Ayed, Information maximization for few-shot learning, Advances in Neural Information Processing Systems 33 (2020) 2445–2457
2020
-
[33]
Zhang, C
J. Zhang, C. Zhao, B. Ni, M. Xu, X. Yang, Variational few-shot learning, in: Proceedings of the IEEE /CVF In- ternational Conference on Computer Vision, 2019, pp. 1685–1694
2019
-
[34]
L. Qiao, Y . Shi, J. Li, Y . Wang, T. Huang, Y . Tian, Trans- ductive episodic-wise adaptive metric for few-shot learn- ing, in: Proceedings of the IEEE/CVF international con- ference on computer vision, 2019, pp. 3603–3612
2019
-
[35]
Hajimiri, M
S. Hajimiri, M. Boudiaf, I. Ben Ayed, J. Dolz, A strong baseline for generalized few-shot semantic segmenta- tion, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 11269–11278
2023
-
[36]
Gavves, T
E. Gavves, T. Mensink, T. Tommasi, C. G. Snoek, T. Tuytelaars, Active transfer learning with zero-shot pri- ors: Reusing past datasets for future tasks, in: Proceed- ings of the IEEE International Conference on Computer Vision, 2015, pp. 2731–2739
2015
-
[37]
Mosbach, T
M. Mosbach, T. Pimentel, S. Ravfogel, D. Klakow, Y . Elazar, Few-shot fine-tuning vs. in-context learn- ing: A fair comparison and evaluation, arXiv preprint arXiv:2305.16938 (2023)
2023 arXiv
-
[38]
Eustratiadis, Ł
P. Eustratiadis, Ł. Dudziak, D. Li, T. Hospedales, Neural fine-tuning search for few-shot learning, arXiv preprint arXiv:2306.09295 (2023)
2023 arXiv
-
[39]
P. Peng, J. Wang, How to fine-tune deep neural networks in few-shot learning?, arXiv preprint arXiv:2012.00204 (2020)
2020 arXiv
-
[40]
H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, C. A. Ra ffel, Few-shot parameter-e fficient fine-tuning is better and cheaper than in-context learning, Advances in Neural Information Processing Systems 35 (2022) 1950–1965
2022
-
[41]
Z. Shen, Z. Liu, J. Qin, M. Savvides, K.-T. Cheng, Partial is better than all: revisiting fine-tuning strategy for few- shot learning, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 35, 2021, pp. 9594–9602
2021
-
[42]
S. X. Hu, D. Li, J. Stühmer, M. Kim, T. M. Hospedales, Pushing the limits of simple pipelines for few-shot learn- ing: External data and fine-tuning make a di fference, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 9068–9077
2022
-
[43]
Y . Guo, H. Shi, A. Kumar, K. Grauman, T. Rosing, R. Feris, Spottune: transfer learning through adaptive fine-tuning, in: Proceedings of the IEEE /CVF confer- ence on computer vision and pattern recognition, 2019, pp. 4805–4814
2019
-
[44]
C. Li, S. Li, H. Wang, F. Gu, A. D. Ball, Attention-based deep meta-transfer learning for few-shot fine-grained fault diagnosis, Knowledge-Based Systems 264 (2023) 110345
2023
-
[45]
X. Xu, Z. Wang, Z. Chi, H. Yang, W. Du, Complemen- tary features based prototype self-updating for few-shot learning, Expert Systems with Applications 214 (2023) 119067
2023
-
[46]
Shafahi, P
A. Shafahi, P. Saadatpanah, C. Zhu, A. Ghiasi, C. Studer, D. Jacobs, T. Goldstein, Adversarially robust transfer learning, arXiv preprint arXiv:1905.08232 (2019). 21
2019 arXiv
-
[47]
E. Real, C. Liang, D. So, Q. Le, Automl-zero: Evolving machine learning algorithms from scratch, in: Interna- tional conference on machine learning, PMLR, 2020, pp. 8007–8019
2020
-
[48]
Abuduweili, X
A. Abuduweili, X. Li, H. Shi, C.-Z. Xu, D. Dou, Adaptive consistency regularization for semi-supervised transfer learning, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2021, pp. 6923–6932
2021
-
[49]
D. Yan, J. Huang, H. Sun, F. Ding, Few-shot object de- tection with weight imprinting, Cognitive Computation (2023) 1–11
2023
-
[50]
Hospedales, A
T. Hospedales, A. Antoniou, P. Micaelli, A. Storkey, Meta-learning in neural networks: A survey, IEEE trans- actions on pattern analysis and machine intelligence 44 (9) (2021) 5149–5169
2021
-
[51]
Vinyals, C
O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra, et al., Matching networks for one shot learning, Advances in neural information processing systems 29 (2016)
2016
-
[52]
Antoniou, A
A. Antoniou, A. Storkey, H. Edwards, Data augmen- tation generative adversarial networks, arXiv preprint arXiv:1711.04340 (2017)
2017 arXiv
-
[53]
Shorten, T
C. Shorten, T. M. Khoshgoftaar, A survey on image data augmentation for deep learning, Journal of Big Data 6 (1) (2019) 69
2019
-
[54]
Lemley, S
J. Lemley, S. Bazrafkan, P. Corcoran, Smart augmenta- tion learning an optimal data augmentation strategy, Ieee Access 5 (2017) 5858–5869
2017
-
[55]
Goodfellow, Y
I. Goodfellow, Y . Bengio, A. Courville, Regularization for deep learning, Deep learning (2016) 216–261
2016
-
[56]
Srivastava, G
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, R. Salakhutdinov, Dropout: a simple way to prevent neu- ral networks from overfitting, The journal of machine learning research 15 (1) (2014) 1929–1958
2014
-
[57]
Kuka ˇcka, V
J. Kuka ˇcka, V . Golkov, D. Cremers, Regulariza- tion for deep learning: A taxonomy, arXiv preprint arXiv:1710.10686 (2017)
2017 arXiv
-
[58]
Y . Li, P. Zhang, X. Xu, Y . Lai, F. Shen, L. Chen, P. Gao, Few-shot prototype alignment regularization network for document image layout segementation, Pattern Recogni- tion 115 (2021) 107882
2021
-
[59]
Girshick, J
R. Girshick, J. Donahue, T. Darrell, J. Malik, Rich fea- ture hierarchies for accurate object detection and seman- tic segmentation, in: Proceedings of the IEEE confer- ence on computer vision and pattern recognition, 2014, pp. 580–587
2014
-
[60]
Cheng, X
G. Cheng, X. Yuan, X. Yao, K. Yan, Q. Zeng, X. Xie, J. Han, Towards large-scale small object detection: Sur- vey and benchmarks, IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)
2023
-
[61]
Girshick, Fast r-cnn, in: Proceedings of the IEEE international conference on computer vision, 2015, pp
R. Girshick, Fast r-cnn, in: Proceedings of the IEEE international conference on computer vision, 2015, pp. 1440–1448
2015
-
[62]
S. Ren, K. He, R. Girshick, J. Sun, Faster r-cnn: Towards real-time object detection with region proposal networks, Advances in neural information processing systems 28 (2015)
2015
-
[63]
Redmon, S
J. Redmon, S. Divvala, R. Girshick, A. Farhadi, You only look once: Unified, real-time object detection, in: Pro- ceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 779–788
2016
-
[64]
Redmon, A
J. Redmon, A. Farhadi, Yolo9000: better, faster, stronger, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7263–7271
2017
-
[65]
Redmon, A
J. Redmon, A. Farhadi, Yolov3: An incremental im- provement, arXiv preprint arXiv:1804.02767 (2018)
2018 arXiv
-
[67]
YOLOv5 by Ultralytics, https://github.com/ultra lytics/yolov5, accessed: 2023-08-31
2023
-
[68]
C. Li, L. Li, H. Jiang, K. Weng, Y . Geng, L. Li, Z. Ke, Q. Li, M. Cheng, W. Nie, et al., Yolov6: A single-stage object detection framework for industrial applications, arXiv preprint arXiv:2209.02976 (2022)
2022 arXiv
-
[69]
C.-Y . Wang, A. Bochkovskiy, H.-Y . M. Liao, Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 7464–7475
2023
-
[70]
YOLOv8 by Ultralytics, https://github.com/ultra lytics/ultralytics, accessed: 2023-08-31
2023
-
[71]
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.- Y . Fu, A. C. Berg, Ssd: Single shot multibox detector, in: Computer Vision–ECCV 2016: 14th European Con- ference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14, Springer, 2016, pp. 21–37
2016
-
[72]
F. He, N. Gao, Q. Li, S. Du, X. Zhao, K. Huang, Tempo- ral context enhanced feature aggregation for video object detection, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 34, 2020, pp. 10941–10948. 22
2020
-
[73]
S. Lin, F. Qin, H. Peng, R. A. Bly, K. S. Moe, B. Han- naford, Multi-frame feature aggregation for real-time instrument segmentation in endoscopic video, IEEE Robotics and Automation Letters 6 (4) (2021) 6773– 6780
2021
-
[74]
Cores, V
D. Cores, V . M. Brea, M. Mucientes, Spatiotemporal tubelet feature aggregation and object linking for small object detection in videos, Applied Intelligence 53 (1) (2023) 1205–1217
2023
-
[75]
X. Zhu, Y . Wang, J. Dai, L. Yuan, Y . Wei, Flow-guided feature aggregation for video object detection, in: Pro- ceedings of the IEEE international conference on com- puter vision, 2017, pp. 408–417
2017
-
[76]
G. Sun, Y . Liu, H. Ding, T. Probst, L. Van Gool, Coarse- to-fine feature mining for video semantic segmentation, in: Proceedings of the IEEE /CVF Conference on Com- puter Vision and Pattern Recognition, 2022, pp. 3126– 3137
2022
-
[77]
C. Xu, J. Zhang, M. Wang, G. Tian, Y . Liu, Multilevel spatial-temporal feature aggregation for video object de- tection, IEEE Transactions on Circuits and Systems for Video Technology 32 (11) (2022) 7809–7820
2022
-
[78]
Honari, J
S. Honari, J. Yosinski, P. Vincent, C. Pal, Recombinator networks: Learning coarse-to-fine feature aggregation, in: Proceedings of the IEEE conference on computer vi- sion and pattern recognition, 2016, pp. 5743–5752
2016
-
[79]
J. Guo, W. Liu, S. Xin, Z. Zhao, B. Zhang, A frame level feature aggregation method for video target detection, in: 2021 33rd Chinese Control and Decision Conference (CCDC), IEEE, 2021, pp. 1368–1373
2021
-
[80]
Muralidhara, K
S. Muralidhara, K. A. Hashmi, A. Pagani, M. Liwicki, D. Stricker, M. Z. Afzal, Attention-guided disentangled feature aggregation for video object detection, Sensors 22 (21) (2022) 8583
2022
-
[81]
L. Han, P. Wang, Z. Yin, F. Wang, H. Li, Exploiting better feature aggregation for video object detection, in: Proceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 1469–1477
2020
-
[82]
Y . Qian, L. Yu, W. Liu, G. Kang, A. G. Hauptmann, Adaptive feature aggregation for video object detection, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, 2020, pp. 143–147
2020
-
[83]
Q. Zhou, X. Li, L. He, Y . Yang, G. Cheng, Y . Tong, L. Ma, D. Tao, Transvod: end-to-end video object de- tection with spatial-temporal transformers, IEEE Trans- actions on Pattern Analysis and Machine Intelligence (2022)
2022
-
[84]
Carion, F
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kir- illov, S. Zagoruyko, End-to-end object detection with transformers, in: European conference on computer vi- sion, Springer, 2020, pp. 213–229
2020
-
[85]
Y . Wang, Z. Xu, X. Wang, C. Shen, B. Cheng, H. Shen, H. Xia, End-to-end video instance segmentation with transformers, in: Proceedings of the IEEE /CVF confer- ence on computer vision and pattern recognition, 2021, pp. 8741–8750
2021
-
[86]
L. He, Q. Zhou, X. Li, L. Niu, G. Cheng, X. Li, W. Liu, Y . Tong, L. Ma, L. Zhang, End-to-end video object de- tection with spatial-temporal transformers, in: Proceed- ings of the 29th ACM International Conference on Mul- timedia, 2021, pp. 1507–1516
2021
-
[87]
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, B. Guo, Swin transformer: Hierarchical vision trans- former using shifted windows, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012–10022
2021
-
[88]
Cui, Feature aggregated queries for transformer-based video object detectors, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2023, pp
Y . Cui, Feature aggregated queries for transformer-based video object detectors, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2023, pp. 6365–6376
2023
-
[89]
Cui, Faq: Feature aggregated queries for transformer-based video object detectors, arXiv preprint arXiv:2303.08319 (2023)
Y . Cui, Faq: Feature aggregated queries for transformer-based video object detectors, arXiv preprint arXiv:2303.08319 (2023)
2023 arXiv
-
[90]
Z. Gao, Q. Wang, Z. Pan, Z. Zhai, H. Long, Pointpaint- ing: 3d object detection aided by semantic image infor- mation, Sensors 23 (5) (2023) 2868
2023
-
[91]
S. Shi, X. Wang, H. Li, Pointrcnn: 3d object proposal generation and detection from point cloud, in: Proceed- ings of the IEEE /CVF conference on computer vision and pattern recognition, 2019, pp. 770–779
2019
-
[92]
S. Shi, Z. Wang, X. Wang, H. Li, Part-a^ 2 net: 3d part- aware and aggregation neural network for object detec- tion from point cloud, arXiv preprint arXiv:1907.03670 2 (3) (2019)
2019 arXiv
-
[93]
Shreyas, M
E. Shreyas, M. H. Sheth, et al., 3d object detection and tracking methods using deep learning for computer vi- sion applications, in: 2021 International Conference on Recent Trends on Electronics, Information, Communica- tion & Technology (RTEICT), IEEE, 2021, pp. 735–738
2021
-
[94]
A. H. Lang, S. V ora, H. Caesar, L. Zhou, J. Yang, O. Bei- jbom, Pointpillars: Fast encoders for object detection from point clouds, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2019, pp. 12697–12705
2019
-
[95]
Z. Tian, X. Chu, X. Wang, X. Wei, C. Shen, Fully convo- lutional one-stage 3d object detection on lidar range im- ages, Advances in Neural Information Processing Sys- tems 35 (2022) 34899–34911. 23
2022
-
[96]
H. Zhao, M. Tian, S. Sun, J. Shao, J. Yan, S. Yi, X. Wang, X. Tang, Spindle net: Person re-identification with hu- man body region guided feature decomposition and fu- sion, in: Proceedings of the IEEE conference on com- puter vision and pattern recognition, 2017, pp. 1077– 1085
2017
-
[97]
T. Yin, X. Zhou, P. Krahenbuhl, Center-based 3d object detection and tracking, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 11784–11793
2021
-
[98]
H.-S. Kim, M. Won Lee, 3d object recognition using x3d and deep learning, in: The 25th International Conference on 3D Web Technology, 2020, pp. 1–8
2020
-
[99]
X. Chen, K. Kundu, Z. Zhang, H. Ma, S. Fidler, R. Urta- sun, Monocular 3d object detection for autonomous driv- ing, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2147–2156
2016
-
[100]
T. He, S. Soatto, Mono3d ++: Monocular 3d vehicle de- tection with two-scale 3d hypotheses and task priors, in: Proceedings of the AAAI Conference on Artificial Intel- ligence, V ol. 33, 2019, pp. 8409–8416
2019
-
[101]
Y . You, Y . Wang, W.-L. Chao, D. Garg, G. Pleiss, B. Hariharan, M. Campbell, K. Q. Weinberger, Pseudo- lidar++: Accurate depth for 3d object detection in autonomous driving, arXiv preprint arXiv:1906.06310 (2019)
2019 arXiv
-
[102]
H. Liu, C. Wu, H. Wang, Real time object detection using lidar and camera fusion for autonomous driving, Scien- tific Reports 13 (1) (2023) 8056
2023
-
[103]
J. Ku, M. Mozifian, J. Lee, A. Harakeh, S. L. Waslan- der, Joint 3d proposal generation and object detection from view aggregation, in: 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2018, pp. 1–8
2018
-
[104]
V ora, A
S. V ora, A. H. Lang, B. Helou, O. Beijbom, Pointpaint- ing: Sequential fusion for 3d object detection, in: Pro- ceedings of the IEEE /CVF conference on computer vi- sion and pattern recognition, 2020, pp. 4604–4612
2020
-
[105]
D. Xu, D. Anguelov, A. Jain, Pointfusion: Deep sensor fusion for 3d bounding box estimation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 244–253
2018
-
[106]
Belouadah, A
E. Belouadah, A. Dapogny, K. Bailly, Multiod: Rehearsal-free multihead incremental object detector, arXiv preprint arXiv:2309.05334 (2023)
2023 arXiv
-
[107]
Krothapalli, L
U. Krothapalli, L. Abbott, One size doesn’t fit all: Adap- tive label smoothing (2020)
2020
-
[108]
W. Lv, S. Xu, Y . Zhao, G. Wang, J. Wei, C. Cui, Y . Du, Q. Dang, Y . Liu, Detrs beat yolos on real-time object detection, arXiv preprint arXiv:2304.08069 (2023)
2023 arXiv
-
[109]
H. Su, Y . He, R. Jiang, J. Zhang, W. Zou, B. Fan, Dsla: Dynamic smooth label assignment for e fficient anchor- free object detection, Pattern Recognition 131 (2022) 108868
2022
-
[110]
T. Liu, L. Zhang, Y . Wang, J. Guan, Y . Fu, J. Zhao, S. Zhou, Recent few-shot object detection algorithms: A survey with performance comparison, ACM Trans- actions on Intelligent Systems and Technology 14 (4) (2023) 1–36
2023
-
[111]
W. Jin, F. Guo, L. Zhu, Incremental self-supervised learning based on transformer for anomaly detection and localization, arXiv preprint arXiv:2303.17354 (2023)
2023 arXiv
-
[112]
Jiang, Z
X. Jiang, Z. Li, M. Tian, J. Liu, S. Yi, D. Miao, Few-shot object detection via improved classification features, in: Proceedings of the IEEE /CVF Winter Conference on Applications of Computer Vision, 2023, pp. 5386–5395
2023
-
[113]
Köhler, M
M. Köhler, M. Eisenbach, H.-M. Gross, Few-shot object detection: a comprehensive survey, IEEE Transactions on Neural Networks and Learning Systems (2023)
2023
-
[114]
A. Wu, S. Zhao, C. Deng, W. Liu, Generalized and dis- criminative few-shot object detection via svd-dictionary enhancement, Advances in Neural Information Process- ing Systems 34 (2021) 6353–6364
2021
-
[115]
B. Kang, Z. Liu, X. Wang, F. Yu, J. Feng, T. Darrell, Few-shot object detection via feature reweighting, in: Proceedings of the IEEE /CVF International Conference on Computer Vision, 2019, pp. 8420–8429
2019
-
[116]
Shangguan, M
Z. Shangguan, M. Rostami, Identification of novel classes for improving few-shot object detection, arXiv preprint arXiv:2303.10422 (2023)
2023 arXiv
-
[117]
X. Wang, T. E. Huang, T. Darrell, J. E. Gonzalez, F. Yu, Frustratingly simple few-shot object detection, arXiv preprint arXiv:2003.06957 (2020)
2020 arXiv
-
[118]
J. Wang, D. Chen, Few-shot object detection method based on knowledge reasoning, Electronics 11 (9) (2022) 1327
2022
-
[120]
Y . Chen, Y . Cao, H. Hu, L. Wang, Memory enhanced global-local aggregation for video object detection, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2020, pp. 10337–10346
2020
-
[121]
K. Lee, H. Yang, S. Chakraborty, Z. Cai, G. Swami- nathan, A. Ravichandran, O. Dabeer, Rethinking few- shot object detection on a multi-domain benchmark, in: European Conference on Computer Vision, Springer, 2022, pp. 366–382. 24
2022
-
[122]
Müller, S
R. Müller, S. Kornblith, G. E. Hinton, When does label smoothing help?, Advances in neural information pro- cessing systems 32 (2019)
2019
-
[123]
G. Han, Y . He, S. Huang, J. Ma, S.-F. Chang, Query adaptive few-shot object detection with heterogeneous graph convolutional networks, in: Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, 2021, pp. 3263–3272
2021
-
[124]
Niklaus, L
S. Niklaus, L. Mai, F. Liu, Video frame interpolation via adaptive separable convolution, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 261–270
2017
-
[125]
M. Xu, S. Yoon, A. Fuentes, D. S. Park, A comprehen- sive survey of image augmentation techniques for deep learning, Pattern Recognition (2023) 109347
2023
-
[126]
H. Wu, C. Song, S. Yue, Z. Wang, J. Xiao, Y . Liu, Dy- namic video mix-up for cross-domain action recognition, Neurocomputing 471 (2022) 358–368
2022
-
[127]
Mangla, N
P. Mangla, N. Kumari, A. Sinha, M. Singh, B. Krishna- murthy, V . N. Balasubramanian, Charting the right mani- fold: Manifold mixup for few-shot learning, in: Proceed- ings of the IEEE/CVF winter conference on applications of computer vision, 2020, pp. 2218–2227
2020
-
[128]
A. Roy, A. Shah, K. Shah, P. Dhar, A. Cherian, R. Chel- lappa, Felmi: few shot learning with hard mixup, Ad- vances in Neural Information Processing Systems 35 (2022) 24474–24486
2022
-
[129]
Nakamura, Y
Y . Nakamura, Y . Ishii, Y . Maruyama, T. Yamashita, Few- shot adaptive object detection with cross-domain cutmix, in: Proceedings of the Asian Conference on Computer Vision, 2022, pp. 1350–1367
2022
-
[130]
J. Yoo, N. Ahn, K.-A. Sohn, Rethinking data augmenta- tion for image super-resolution: A comprehensive analy- sis and a new strategy, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2020, pp. 8375–8384
2020
-
[131]
Zhang, T
C. Zhang, T. Yang, J. Weng, M. Cao, J. Wang, Y . Zou, Unsupervised pre-training for temporal action localiza- tion tasks, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 14031–14041
2022
-
[132]
Aich, K.-C
A. Aich, K.-C. Peng, A. K. Roy-Chowdhury, Cross- domain video anomaly detection without target domain adaptation, in: Proceedings of the IEEE /CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2579–2591
2023
-
[133]
Olsson, W
V . Olsson, W. Tranheden, J. Pinto, L. Svensson, Class- mix: Segmentation-based data augmentation for semi- supervised learning, in: Proceedings of the IEEE /CVF Winter Conference on Applications of Computer Vision, 2021, pp. 1369–1378
2021
-
[134]
A. S. Chakravarthy, W.-D. Jang, Z. Lin, D. Wei, S. Bai, H. Pfister, Object propagation via inter-frame attentions for temporally stable video instance segmentation, arXiv preprint arXiv:2111.07529 (2021)
2021 arXiv
-
[136]
C. Deng, D. Chen, Q. Wu, Identity-consistent ag- gregation for video object detection, arXiv preprint arXiv:2308.07737 (2023)
2023 arXiv
-
[137]
Zhang, H
Y . Zhang, H. Wang, H. Zhu, Z. Chen, Optical flow reusing for high-e fficiency space-time video super res- olution, IEEE Transactions on Circuits and Systems for Video Technology (2022)
2022
-
[138]
J. Lin, X. Hu, Y . Cai, H. Wang, Y . Yan, X. Zou, Y . Zhang, L. Van Gool, Unsupervised flow-aligned sequence-to- sequence learning for video restoration, in: International Conference on Machine Learning, PMLR, 2022, pp. 13394–13404
2022
-
[139]
X. Du, Y . Li, Y . Cui, R. Qian, J. Li, I. Bello, Revis- iting 3d resnets for video recognition, arXiv preprint arXiv:2109.01696 (2021)
2021 arXiv
-
[140]
Z. Ma, H. Zhang, J. Liu, Ms-lstm: Exploring spatiotem- poral multiscale representations in video prediction do- main, arXiv preprint arXiv:2304.07724 (2023)
2023 arXiv
-
[141]
P. L. Jeune, A. Mokraoui, A unified framework for attention-based few-shot object detection, arXiv preprint arXiv:2201.02052 (2022)
2022 arXiv
-
[142]
P. Pal, P. Chattopadhyay, M. Swarnkar, Temporal fea- ture aggregation with attention for insider threat detec- tion from activity logs, Expert Systems with Applica- tions 224 (2023) 119925
2023
-
[143]
Gordevi ˇcius, J
J. Gordevi ˇcius, J. Gamper, M. Böhlen, Parsimonious temporal aggregation, in: Proceedings of the 12th In- ternational Conference on Extending Database Technol- ogy: Advances in Database Technology, 2009, pp. 1006– 1017
2009
-
[144]
Y . Fu, S. Sen, J. Reimann, C. Theurer, Spatiotempo- ral representation learning with gan trained lstm-lstm networks, in: 2020 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 2020, pp. 10548–10555
2020
-
[145]
S. Dai, Y . Yu, H. Fan, J. Dong, Spatio-temporal represen- tation learning with social tie for personalized poi recom- mendation, Data Science and Engineering 7 (1) (2022) 44–56
2022
-
[146]
Jin, Y .-F
M. Jin, Y .-F. Li, Y . Zheng, B. Yang, S. Pan, Spatiotempo- ral representation learning on time series with dynamic graph odes (2021). 25
2021
-
[147]
F. He, Q. Li, X. Zhao, K. Huang, Temporal-adaptive sparse feature aggregation for video object detection, Pattern Recognition 127 (2022) 108587
2022
-
[148]
Nirthika, S
R. Nirthika, S. Manivannan, A. Ramanan, R. Wang, Pooling in convolutional neural networks for medical image analysis: a survey and an empirical study, Neural Computing and Applications 34 (7) (2022) 5321–5347
2022
-
[149]
X. Luo, X. Tu, Y . Ding, G. Gao, M. Deng, Expec- tation pooling: an e ffective and interpretable pooling method for predicting dna–protein binding, Bioinformat- ics 36 (5) (2020) 1405–1412
2020
-
[150]
Rouvier, P.-M
M. Rouvier, P.-M. Bousquet, J. Duret, Study on the tem- poral pooling used in deep neural networks for speaker verification, in: 2021 29th European Signal Processing Conference (EUSIPCO), IEEE, 2021, pp. 501–505
2021
-
[151]
Zafar, M
A. Zafar, M. Aamir, N. Mohd Nawi, A. Arshad, S. Riaz, A. Alruban, A. K. Dutta, S. Almotairi, A comparison of pooling methods for convolutional neural networks, Applied Sciences 12 (17) (2022) 8643
2022
-
[152]
A.-K. N. Vu, N.-D. Nguyen, K.-D. Nguyen, V .-T. Nguyen, T. D. Ngo, T.-T. Do, T. V . Nguyen, Few-shot object detection via baby learning, Image and Vision Computing 120 (2022) 104398
2022
-
[153]
Kim, H.-G
G. Kim, H.-G. Jung, S.-W. Lee, Spatial reasoning for few-shot object detection, Pattern Recognition 120 (2021) 108118
2021
-
[154]
S. Zhao, X. Qi, Prototypical votenet for few-shot 3d point cloud object detection, Advances in Neural Infor- mation Processing Systems 35 (2022) 13838–13851
2022
-
[155]
J. Liu, X. Dong, S. Zhao, J. Shen, Generalized few-shot 3d object detection of lidar point cloud for autonomous driving, arXiv preprint arXiv:2302.03914 (2023)
2023 arXiv
-
[157]
M. Guo, E. Chou, D.-A. Huang, S. Song, S. Yeung, L. Fei-Fei, Neural graph matching networks for fewshot 3d action recognition, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 653– 669
2018
-
[160]
F. Liu, S. Yang, D. Chen, H. Huang, J. Zhou, Few-shot classification guided by generalization error bound, Pat- tern Recognition (2023) 109904
2023
-
[161]
J. He, Y . Chen, N. Wang, Z. Zhang, 3d video object de- tection with learnable object-centric global optimization, in: Proceedings of the IEEE /CVF Conference on Com- puter Vision and Pattern Recognition, 2023, pp. 5106– 5115
2023
-
[162]
Brazil, G
G. Brazil, G. Pons-Moll, X. Liu, B. Schiele, Kinematic 3d object detection in monocular video, in: Computer Vision–ECCV 2020: 16th European Conference, Glas- gow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, Springer, 2020, pp. 135–152
2020
-
[163]
J. Wu, L. Song, T. Wang, Q. Zhang, J. Yuan, Forest r- cnn: Large-vocabulary long-tailed object detection and instance segmentation, in: Proceedings of the 28th ACM international conference on multimedia, 2020, pp. 1570– 1578
2020
-
[164]
J. Wu, S. Liu, D. Huang, Y . Wang, Multi-scale posi- tive sample refinement for few-shot object detection, in: Computer Vision–ECCV 2020: 16th European Confer- ence, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVI 16, Springer, 2020, pp. 456–472
2020
-
[165]
D. Lee, J. Kim, Resolving class imbalance for lidar- based object detector by dynamic weight average and contextual ground truth sampling, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Com- puter Vision, 2023, pp. 682–691
2023
-
[166]
Z. Yu, G. Wang, L. Chen, S. Raschka, J. Luo, Few-shot learning for video object detection in a transfer-learning scheme, arXiv preprint arXiv:2103.14724 (2021)
2021 arXiv
-
[167]
S. Baik, M. Choi, J. Choi, H. Kim, K. M. Lee, Meta- learning with adaptive hyperparameters, Advances in neural information processing systems 33 (2020) 20755– 20765
2020
-
[168]
Rajendran, A
J. Rajendran, A. Irpan, E. Jang, Meta-learning requires meta-augmentation, Advances in Neural Information Processing Systems 33 (2020) 5705–5715
2020
-
[169]
C. Si, X. Nie, W. Wang, L. Wang, T. Tan, J. Feng, Ad- versarial self-supervised learning for semi-supervised 3d action recognition, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part VII 16, Springer, 2020, pp. 35–51
2020
-
[170]
G. Yang, D. Sun, V . Jampani, D. Vlasic, F. Cole, C. Liu, D. Ramanan, Viser: Video-specific surface em- beddings for articulated 3d shape reconstruction, Ad- vances in Neural Information Processing Systems 34 (2021) 19326–19338. 26
2021
-
[171]
Huang, I
G. Huang, I. Laradji, D. Vazquez, S. Lacoste-Julien, P. Rodriguez, A survey of self-supervised and few-shot object detection, IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (4) (2022) 4071–4089
2022
-
[172]
F. Bao, G. Wu, C. Li, J. Zhu, B. Zhang, Stability and generalization of bilevel programming in hyperparam- eter optimization, Advances in neural information pro- cessing systems 34 (2021) 4529–4541
2021
-
[173]
X. Luo, H. Wu, J. Zhang, L. Gao, J. Xu, J. Song, A closer look at few-shot classification again, arXiv preprint arXiv:2301.12246 (2023)
2023 arXiv
-
[174]
Zimmer, M
W. Zimmer, M. Grabler, A. Knoll, Real-time and robust 3d object detection within road-side lidars using domain adaptation, arXiv preprint arXiv:2204.00132 (2022)
2022 arXiv
-
[175]
Y . Wang, J. Yin, W. Li, P. Frossard, R. Yang, J. Shen, Ssda3d: Semi-supervised domain adaptation for 3d ob- ject detection from point cloud, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 37, 2023, pp. 2707–2715
2023
-
[176]
J. Yang, S. Shi, Z. Wang, H. Li, X. Qi, St3d: Self- training for unsupervised domain adaptation on 3d object detection, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2021, pp. 10368–10378
2021
-
[177]
Hegde, V
D. Hegde, V . Kilic, V . Sindagi, A. B. Cooper, M. Foster, V . M. Patel, Source-free unsupervised domain adaptation for 3d object detection in adverse weather, in: 2023 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 2023, pp. 6973–6980
2023
-
[179]
J. Han, Y . Ren, J. Ding, K. Yan, G.-S. Xia, Few-shot ob- ject detection via variational feature aggregation, arXiv preprint arXiv:2301.13411 (2023)
2023 arXiv
-
[180]
I. H. Sarker, Machine learning: Algorithms, real-world applications and research directions, SN computer sci- ence 2 (3) (2021) 160
2021
-
[181]
S. S. A. Zaidi, M. S. Ansari, A. Aslam, N. Kanwal, M. Asghar, B. Lee, A survey of modern deep learning based object detection models, Digital Signal Processing 126 (2022) 103514
2022
-
[182]
H. Wang, X. Zhang, Y . Hu, Y . Yang, X. Cao, X. Zhen, Few-shot semantic segmentation with democratic atten- tion networks, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIII 16, Springer, 2020, pp. 730–746
2020
-
[183]
C. Guo, B. Fan, Q. Zhang, S. Xiang, C. Pan, Augfpn: Im- proving multi-scale feature learning for object detection, in: Proceedings of the IEEE /CVF conference on com- puter vision and pattern recognition, 2020, pp. 12595– 12604
2020
-
[184]
R. Cao, K. Zhang, Y . Chen, X. Yang, C. Jin, Point cloud completion via multi-scale edge convolution and atten- tion, in: Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 6183–6192
2022
-
[185]
Z. Li, P. Xu, X. Chang, L. Yang, Y . Zhang, L. Yao, X. Chen, When object detection meets knowledge distil- lation: A survey, IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)
2023
-
[186]
X. He, K. Zhao, X. Chu, Automl: A survey of the state-of-the-art, Knowledge-Based Systems 212 (2021) 106622
2021
-
[187]
C. Chen, J. Wang, J. Pan, C. Bian, Z. Zhang, Graphskt: graph-guided structured knowledge transfer for domain adaptive lesion detection, IEEE Transactions on Medical Imaging 42 (2) (2022) 507–518
2022
-
[188]
K. Liu, S. Lyu, P. Shivakumara, Y . Lu, Few-shot ob- ject segmentation with a new feature aggregation mod- ule, Displays 78 (2023) 102459
2023
-
[189]
L. Zhao, G. Liu, D. Guo, W. Li, X. Fang, Boosting few- shot visual recognition via saliency-guided complemen- tary attention, Neurocomputing 507 (2022) 412–427
2022
-
[190]
Munjal, A
B. Munjal, A. Flaborea, S. Amin, F. Tombari, F. Galasso, Query-guided networks for few-shot fine-grained clas- sification and person search, Pattern Recognition 133 (2023) 109049
2023
-
[191]
Mahapatra, Interpretable saliency maps and self- supervised learning for generalized zero shot medical image classification, arXiv preprint arXiv:2204.01728 (2022)
D. Mahapatra, Interpretable saliency maps and self- supervised learning for generalized zero shot medical image classification, arXiv preprint arXiv:2204.01728 (2022)
2022 arXiv
-
[192]
W. Wang, L. Duan, Y . Wang, J. Fan, Z. Gong, Z. Zhang, A survey of deep visual cross-domain few-shot learning, arXiv preprint arXiv:2303.09253 (2023)
2023 arXiv
-
[193]
J. Zhu, J. Liu, S. Yang, Q. Zhang, X. He, Open bench- marking for click-through rate prediction, in: Proceed- ings of the 30th ACM International Conference on In- formation & Knowledge Management, 2021, pp. 2759– 2769
2021
-
[194]
S. Sun, Y . Lu, S. Yu, X. Li, Z. Li, Z. Cao, Z. Liu, D. Ye, J. Bao, Rethinking dense retrieval’s few-shot abil- ity, arXiv preprint arXiv:2304.05845 (2023)
2023 arXiv
-
[195]
X. Wang, L. Lian, S. X. Yu, Unsupervised selective la- beling for more e ffective semi-supervised learning, in: European Conference on Computer Vision, Springer, 2022, pp. 427–445. 27
2022
-
[196]
McClurg, A
C. McClurg, A. Ayub, H. Tyagi, S. M. Rajtmajer, A. R. Wagner, Active class selection for few-shot class- incremental learning, arXiv preprint arXiv:2307.02641 (2023)
2023 arXiv
-
[197]
Z. Chen, J. Ge, H. Zhan, S. Huang, D. Wang, Pareto self- supervised training for few-shot learning, in: Proceed- ings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 13663–13672
2021
-
[198]
Perez-Rua, X
J.-M. Perez-Rua, X. Zhu, T. M. Hospedales, T. Xiang, Incremental few-shot object detection, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 13846–13855
2020
-
[199]
Mousavian, D
A. Mousavian, D. Anguelov, J. Flynn, J. Kosecka, 3d bounding box estimation using deep learning and geom- etry, in: Proceedings of the IEEE conference on Com- puter Vision and Pattern Recognition, 2017, pp. 7074– 7082
2017
-
[200]
Manhardt, W
F. Manhardt, W. Kehl, A. Gaidon, Roi-10d: Monocular lifting of 2d detection to 6d pose and metric shape, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 2069–2078
2019
-
[201]
C. R. Qi, W. Liu, C. Wu, H. Su, L. J. Guibas, Frus- tum pointnets for 3d object detection from rgb-d data, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 918–927
2018
-
[202]
X. Bai, Z. Hu, X. Zhu, Q. Huang, Y . Chen, H. Fu, C.-L. Tai, Transfusion: Robust lidar-camera fusion for 3d ob- ject detection with transformers, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 1090–1099
2022
-
[203]
Wang, et al., Data augmentation for meta-learning, arXiv preprint arXiv:2002.08973 (2020)
Y . Wang, et al., Data augmentation for meta-learning, arXiv preprint arXiv:2002.08973 (2020)
2020 arXiv
-
[204]
Franceschi, et al., A theoretical analysis of the num- ber of samples needed to estimate information-theoretic quantities, arXiv preprint arXiv:1708.01974 (2017)
J.-Y . Franceschi, et al., A theoretical analysis of the num- ber of samples needed to estimate information-theoretic quantities, arXiv preprint arXiv:1708.01974 (2017)
2017 arXiv
-
[205]
Few-Shot Learning in Video and 3D Object Detection: A Survey
E. D. Cubuk, et al., Randaugment: Practical data augmentation with no separate search, arXiv preprint arXiv:1909.13719 (2019). 28 The supplementary materials provide additional details and visual overviews to support the survey paper “Few-Shot Learning in Video and 3D Object D...
2019 arXiv
-
[206]
Foundations of Few-Shot Learning : Discusses key concepts like episodic training, problem formulations, meta-learning algorithms, metric-based approaches, data augmentation, and regularization techniques for few-shot learning
-
[207]
Foundations of Object Detection : Provides an overview of object detection methods, including two-stage and one-stage detectors, as well as video and 3D object detection approaches
-
[208]
Few-Shot Video Object Detection: Presents example frameworks and architectures tailored for few-shot video object detec- tion, highlighting techniques like metric learning, temporal feature aggregation, and episodic training
-
[209]
Few-Shot 3D Object Detection : Covers specialized few-shot detection methods for 3D data such as LiDAR point clouds, using techniques like geometric prototypes, support set guidance, and incremental learning. The supplementary materials expand on the key concepts, architecture...
-
[210]
learning to learn
Foundations of Few-Shot Learning Few-shot learning (FSL) has emerged as a critical area of study within the deep learning framework, addressing one of the most pressing challenges in machine learning: the need for vast amounts of labeled data. In many real-world scenarios, obt...
-
[211]
Regions of Interest
Foundations of Object Detection This section begins with an overview of object detection, including common techniques and applications. It then discusses various techniques used in video and 3D object detection. 3.1. Object Detection Object detection (OD) is a cornerstone of c...
-
[212]
and YOLO[63], which apply convolutional filters across an image in a single shot to directly output object locations and classes. The tradeoff between accuracy and speed makes two-stage detectors preferable for applications where accuracy is critical, while one-stage detectors...
-
[213]
Another line of work aggregates points into compact representations like pillars which encode vertical point columns, before applying e fficient 2D con- volutions on pseudo images
extend it for 3D detection by first generating proposals which are then refined using point features. Another line of work aggregates points into compact representations like pillars which encode vertical point columns, before applying e fficient 2D con- volutions on pseudo im...
-
[214]
It first generates initial boxes from LiDAR, then fuses image features in the second decoder layer using soft-attention
method fuses LiDAR and images using a novel transformer architecture with soft-attention. It first generates initial boxes from LiDAR, then fuses image features in the second decoder layer using soft-attention. This provides robustness to misalignment 34 and degraded image qua...
-
[215]
Few-Shot Video Object Detection 37 Figure S4: Overview of the Prototypical V oteNet architecture for few-shot 3D object detection [154]. It contains two key components - the Prototypical V ote Module (PVM) which refines local features using geometric prototypes, and the Protot...
-
[216]
It consists of a 3D Meta-Detector module that generates class-specific reweighting vectors zn from the few-shot support points
Few-Shot 3D Object Detection 39 Figure S7: Overview of the MetaDet3D framework for few-shot 3D object detection [156]. It consists of a 3D Meta-Detector module that generates class-specific reweighting vectors zn from the few-shot support points. These reweighting vectors guid...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.