REVIEW 3 major objections 5 minor 42 references
Multi-Point Positional Insertion Tuning for Small Object Detection
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Adding a small set of learnable positional embeddings to a frozen object detector matches prompt-tuning performance for small objects while using about 0.5 million learnable parameters.
desk verdict The empirical parity claim at 0.5M parameters holds on the reported numbers, but the 'positional' mechanism is really an input-independent bias, and the paper needs a bias-only control before that interpretation is taken seriously. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the multi-head positional encoder (MHP encoder). It takes sinusoidal positional embeddings (with $D=64$, $L=80{,}000$, $C=10{,}000$), passes them through $M$ tiny MLPs of two linear-LayerNorm-SwiGLU blocks, and then combines the $M$ streams into $N=26$ embeddings via a learnable multi-head mixer with weights $A_{ij}$. Each result is shaped by a linear layer and added to a latent feature as $h'_i(x) = h_i(x) + p_i$. The multi-point placement is what connects the parameter budget to the architecture: two points after the BERT and Swin encoders, two per feature-enhancer block (twelve total), and two per decoder block (twelve total). The mixer is what keeps the parameter count low, letting $M<N$ streams share the $N$ insertion positions.
What would settle it
Run the identical MPI protocol on a second small-object benchmark (or on a sampled subset of a general detection dataset) with the same 26 insertion points and 0.50M parameter budget, and check whether the validation mAP gap relative to CoOp with a decoder stays within about one point. Alternatively, replace the sinusoidal embeddings with fixed random embeddings of the same shape and see whether the feature-enhancer ablation drop (26.5 to 24.6) disappears, which would show that the positional format, not just the additive perturbation, is what carries the benefit.
Extended reading notes
Core claim
The central discovery is that a frozen vision-language detector can be adapted to a small-object benchmark by adding learned position-dependent vectors to its latent features at selected layers, with no other parameter updates. MPI tuning trains only a multi-head positional encoder whose outputs $p_i$ are added to each selected latent feature $h_i(x)$. On the SODA-D test set this yields 25.7 mAP with 0.50M parameters, effectively matching CoOp w/ dec and VPT w/ dec while using roughly 1/24 of their learnable parameters. The paper also reports that the feature-enhancer insertion points carry most of the benefit: removing them drops validation mAP from 26.5 to 24.6, while removing the decoder insertion points leaves mAP unchanged.
Load-bearing premise
The method assumes that a frozen pretrained detector can be adapted to small-object detection by adding learnable position-dependent vectors to 26 selected latent features, and that this additive form can express the needed feature changes; if that assumption fails on other datasets, the observed parity with prompt tuning may not persist.
Editorial extensions
If this is right
- MPI tuning's parameter count scales with the number of tiny MLPs $M$: halving $M$ from 24 to 12 roughly halves the learnable parameters to 0.50M while losing only 0.2 mAP (26.7 to 26.5), offering a direct accuracy-versus-parameter knob.
- Inserting position at the feature enhancer is the critical decision, since the ablation drops about 1.9 mAP when those points are removed, identifying where positional adaptation matters most in a vision-language detector.
- MPI tuning transfers across image-encoder backbones (Swin-T, Swin-B, Swin-L), beating CoOp without decoder finetuning at the same 0.50M parameter budget.
- Because the pretrained detector stays frozen, the method avoids backpropagation through the full 173M-parameter model during finetuning, cutting memory and compute relative to full finetuning.
Reading between the lines
- A natural next test is whether the positional content itself matters: replacing the sinusoidal inputs with random-but-fixed embeddings of the same shape, while keeping the insertion layout, would isolate whether the gain comes from position information or merely from an additive perturbation.
- The method has been validated on only one street-scene dataset, so whether 26 hand-picked insertion points generalize to other small-object domains remains open; an automatic procedure for selecting insertion depths would be a direct extension.
- Because the mixer linearly combines $M$ learned streams into $N$ positions, the effective correction to the latent trajectory is low-rank; this suggests MPI could be combined with adapter or prompt tuning rather than only competing with them, since it occupies a different, additive subspace.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes multi-point positional insertion (MPI) tuning, a parameter-efficient fine-tuning method for small object detection. The method inserts learnable positional embeddings at 26 selected points of a frozen Grounding DINO model, using a multi-head positional encoder built from sinusoidal embeddings, tiny MLPs, and a linear mixer. The adapted features are computed as h'_i(x) = h_i(x) + p_i, where p_i depends on the sinusoidal position index and learned weights but not on the input. On the SODA-D dataset, MPI tuning reports 25.7 mAP with 0.50M learnable parameters, comparable to CoOp (25.8 mAP, 12.00M) and VPT (25.4 mAP, 11.98M) with decoder fine-tuning, while using far fewer parameters. The paper includes ablations on insertion-point groups, the number of tiny MLPs, and three Swin backbones.
Significance. If the empirical results are reliable, MPI tuning is a useful PEFT variant for small object detection: it reaches parity with prompt-tuning baselines at a substantially lower parameter count and shows consistent behavior across backbones. The paper's strengths are its direct comparison against several PEFT baselines, the inclusion of parameter counts, and ablations on module placement and capacity. However, the central mechanistic claim that the gains come from 'precise positional information' is not yet established, because the inserted term p_i is input-independent and could act as a learned bias. The reported 0.1 mAP difference between MPI and CoOp in Table I is also within the range of typical run-to-run variation, and no error bars, multiple seeds, or significance tests are provided.
major comments (3)
- [Sec. III-C, Eqs. (1)-(3)] The MHP encoder consumes only the sinusoidal positional embedding e and never the input x or the latent feature h_i(x). Consequently, after training, each p_i is a fixed, input-independent tensor at each insertion point, and Eq. (1) is an additive per-point bias. The paper attributes the improvement to 'providing precise positional information to latent features,' but no control experiment replaces the MHP encoder with a directly learned bias tensor at the same 26 points and the same 0.50M parameter budget. Without that control, the results in Table I are equally consistent with a standard bias-tuning PEFT, and the positional interpretation is unsupported.
- [Sec. IV-B, Table I] The main comparison reports a single run per method with no error bars, multiple seeds, or significance tests. The central claim of comparability rests on a 0.1 mAP difference between MPI (25.7) and CoOp with decoder (25.8), which is smaller than typical seed-to-seed variation for detection fine-tuning. At minimum, the authors should report the standard deviation over at least three seeds for the main configurations, or otherwise provide evidence that the observed parity is not due to noise.
- [Sec. III-D and Table II] The choice of the 26 insertion points is described as manually selected, but the paper provides no sensitivity analysis over the insertion-point selection itself. Table II shows that removing the feature-enhancer positions degrades mAP from 26.5 to 24.6, while removing the decoder positions has no effect, but this does not establish that the specific 26 points, or the 12+12+2 split, are preferable to, say, a uniform or random sampling of candidate positions. A comparison with a simpler selection rule would strengthen the 'multi-point' contribution.
minor comments (5)
- [Sec. III-C] There is a typo in the sentence 'Each linear layer is designed to match the shapes of pi and hi(x) ti ensure that pi can be added to hi(x)'; 'ti' should be 'to'.
- [Eq. (2)] The index range for k is written as k = 0, 2, ..., D/2 - 1. Standard sinusoidal embeddings use k = 0, 1, ..., D/2 - 1; the printed formula would skip odd indices and is likely a typo.
- [Sec. III-D and Fig. 1] Figure 1 states that the frozen model is shown with N sequential modules, but the actual 26 insertion points are distributed over components with very different roles (text encoder, image encoder, feature enhancer, decoder). Calling them 'sequential modules' is misleading and should be reworded.
- [Eq. (3) and Sec. III-C] The notation 'Aij ∈ R^{N×M}' is confusing: A_{ij} with subscripts suggests a scalar entry, while the text uses it as a matrix. Use A ∈ R^{N×M} with entries A_{ij}, or an equivalent clearer notation.
- [Sec. IV-B, Table II] The ablation row 'w/o input pos.' is not defined in the text. If it refers to removing the two insertion points after BERT and Swin, the caption or text should say so explicitly.
Circularity Check
No significant circularity: MPI tuning is an empirical comparison against external baselines, with no prediction that reduces to its fitted inputs by construction.
full rationale
This is an empirical methods paper rather than a derivation chain. The central claim is that MPI tuning achieves performance comparable to CoOp and VPT on the SODA-D test set while using fewer learnable parameters. That claim is supported by Table I, which reports test-set mAP against external baselines, and by ablations and backbone studies on the validation set. The learnable parameters of the MHP encoder are optimized on the training split, hyperparameters such as the number of tiny MLPs are selected on the validation split, and the headline results are reported on the held-out test split, so no reported number is constructed to equal a fitted value by definition. Equations (1)-(3) define the proposed adaptation mechanism, but they are architectural definitions rather than a derivation of the empirical mAP, and the experimental outcome is not forced by those definitions. The only self-citation, Ref. [33] by Otake, Kawakami, and Inoue, appears in the related-work discussion as an example of layer adapter tuning and is not load-bearing for the proposed method or its evaluation. The skeptical concern that p_i is input-independent because the MHP encoder consumes only sinusoidal positional embeddings is a legitimate experimental-design critique about missing bias-only controls, but it is not circularity: the paper's comparison against CoOp, VPT, and adapter tuning remains an externally benchmarked empirical result and does not reduce by equation to the fitted parameters. No circular step can be quoted or exhibited, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (5)
- M (number of tiny MLPs) =
12
- N (number of insertion points) =
26
- D (positional embedding dimension) =
64
- L (number of position indices) =
80,000
- C (sinusoidal wavelength constant) =
10,000
assumptions (5)
- domain assumption Adding learned positional embeddings to latent features of a frozen pretrained detector is sufficient to adapt it to small object detection.
- domain assumption The SODA-D dataset and its official splits are correct and representative for evaluating small object detection methods.
- ad hoc to paper The 26 manually selected insertion points cover the locations where positional information is most useful.
- standard math Standard sinusoidal positional embeddings are a useful input representation for learning task-specific positional corrections.
- domain assumption The chosen training hyperparameters (AdamW, lr 1e-4, 12 epochs, batch 16) are equally fair for all compared methods.
Cite this review
Pith. "Pith review of Multi-Point Positional Insertion Tuning for Small Object Detection." pith.science (2026). https://pith.science/paper/TXNZ6OMC
@misc{pith2026241218090,
author = {Pith},
title = {Pith review of: Multi-Point Positional Insertion Tuning for Small Object Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/TXNZ6OMC}},
note = {Machine review of arXiv:2412.18090}
}
read the original abstract
Small object detection aims to localize and classify small objects within images. With recent advances in large-scale vision-language pretraining, finetuning pretrained object detection models has emerged as a promising approach. However, finetuning large models is computationally and memory expensive. To address this issue, this paper introduces multi-point positional insertion (MPI) tuning, a parameter-efficient finetuning (PEFT) method for small object detection. Specifically, MPI incorporates multiple positional embeddings into a frozen pretrained model, enabling the efficient detection of small objects by providing precise positional information to latent features. Through experiments, we demonstrated the effectiveness of the proposed method on the SODA-D dataset. MPI performed comparably to conventional PEFT methods, including CoOp and VPT, while significantly reducing the number of parameters that need to be tuned.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
J. Noh, W. Bae, W. Lee, J. Seo, and G. Kim, “Better to follow, follow to be better: Towards precise supervision of feature super-resolution for small object detection,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2019, pp. 9725–9734
work page 2019
-
[2]
SOD-MTGAN: Small object detection via multi-task generative adversarial network,
Y . Bai, Y . Zhang, M. Ding, and B. Ghanem, “SOD-MTGAN: Small object detection via multi-task generative adversarial network,” in European Conference on Computer Vision (ECCV) , 2018
work page 2018
-
[3]
Small object detection via coarse-to-fine proposal generation and imitation learning,
X. Yuan, G. Cheng, K. Yan, Q. Zeng, and J. Han, “Small object detection via coarse-to-fine proposal generation and imitation learning,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
work page 2023
-
[4]
Robust small-scale pedestrian detec- tion with cued recall via memory learning,
J.-U. Kim, S. Park, and Y . M. Ro, “Robust small-scale pedestrian detec- tion with cued recall via memory learning,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 3030–3039
work page 2021
-
[5]
Self-mimic learning for small-scale pedestrian detection,
J. Wu, C. Zhou, Q. Zhang, M. Yang, and J. Yuan, “Self-mimic learning for small-scale pedestrian detection,” in ACM International Conference on Multimedia (ACMMM) , 2020, pp. 2012–2020
work page 2020
-
[6]
Dynamic local and global context exploration for small object detection,
Z. Zhang, P. Gong, H. Sun, P. Wu, and X. Yang, “Dynamic local and global context exploration for small object detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5
work page 2023
-
[7]
Grounded language-image pre-training,
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y . Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, K.-W. Chang, and J. Gao, “Grounded language-image pre-training,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022
work page 2022
-
[8]
GLIPv2: Unifying localization and vision-language understanding,
H. Zhang, P. Zhang, X. Hu, Y .-C. Chen, L. H. Li, X. Dai, L. Wang, L. Yuan, J.-N. Hwang, and J. Gao, “GLIPv2: Unifying localization and vision-language understanding,” in Annual Conference on Neural Information Processing Systems (NeurIPS) , 2022
work page 2022
Show all 42 references
-
[9]
Grounding dino: Marrying dino with grounded pre-training for open-set object detection,
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al., “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” in European Conference on Computer Vision (ECCV), 2024
2024
-
[10]
An open and comprehensive pipeline for unified object grounding and detection,
X. Zhao, Y . Chen, S. Xu, X. Li, X. Wang, Y . Li, and H. Huang, “An open and comprehensive pipeline for unified object grounding and detection,” arXiv preprint arXiv:2401.02361 , 2024
2024 arXiv
-
[11]
Multiway- adapter: Adapting multimodal large language models for scalable image- text retrieval,
Z. Long, G. Killick, R. McCreadie, and G. A. Camarasa, “Multiway- adapter: Adapting multimodal large language models for scalable image- text retrieval,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024
2024
-
[12]
Automatic design of adapter architectures for enhanced parameter-efficient fine-tuning,
H. Zhou, X. Wan, I. Vuli ´c, and A. Korhonen, “Automatic design of adapter architectures for enhanced parameter-efficient fine-tuning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024
2024
-
[13]
Test-time distribution learning adapter for cross-modal visual reasoning,
Y . Zhang and C. Zhang, “Test-time distribution learning adapter for cross-modal visual reasoning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024
2024
-
[14]
Adapter-based incremental learning for face forgery detection,
C. Gao, Q. Xu, P. Qiao, K. Xu, X. Qian, and Y . Dou, “Adapter-based incremental learning for face forgery detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024
2024
-
[15]
Parameter-efficient transfer learning for nlp,
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in International Conference on Machine Learning (ICML), 2019, pp. 2790–2799
2019
-
[16]
Learning to prompt for vision- language models,
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision- language models,” International Journal of Computer Vision (IJCV) , 2022
2022
-
[17]
Conditional prompt learning for vision-language models,
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Conditional prompt learning for vision-language models,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022, pp. 16816–16825
2022
-
[18]
Visual prompt tuning,
M. Jia, L. Tang, B.-C. Chen, C. Cardie, S. Belongie, B. Hariharan, and S.-N. Lim, “Visual prompt tuning,” in European Conference on Computer Vision (ECCV) , 2022
2022
-
[19]
Visual prompt tuning for weakly supervised phrase grounding,
P. Lin, Z. Yu, M. Lu, F. Feng, R. Li, and X. Wang, “Visual prompt tuning for weakly supervised phrase grounding,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024
2024
-
[20]
Enhanced transfer learning with efficient modeling and adaptive fusion of knowledge via prompt tuning,
M. Xu, Z. Guo, Y . Zeng, and D. Xiong, “Enhanced transfer learning with efficient modeling and adaptive fusion of knowledge via prompt tuning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024
2024
-
[21]
Cophtc: Contrastive learning with prompt tuning for hierarchical text classification,
F. Cai, Z. Zhang, D. Liu, X. Fang, and J. Tong, “Cophtc: Contrastive learning with prompt tuning for hierarchical text classification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024
2024
-
[22]
To- wards large-scale small object detection: Survey and benchmarks,
G. Cheng, X. Yuan, X. Yao, K. Yan, Q. Zeng, X. Xie, and J. Han, “To- wards large-scale small object detection: Survey and benchmarks,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), pp. 1–20, 2023
2023
-
[23]
Focal loss for dense object detection,
T.-Y . Lin, P. Goyal, R. Girshick, K. He, and P. Doll ´ar, “Focal loss for dense object detection,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2017, pp. 2980–2988
2017
-
[24]
FCOS: Fully convolutional one-stage object detection,
Z. Tian, C. Shen, H. Chen, and T. He, “FCOS: Fully convolutional one-stage object detection,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2019, pp. 9627–9636
2019
-
[25]
Faster R-CNN: Towards real-time object detection with region proposal networks,
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in Annual Conference on Neural Information Processing Systems (NeurIPS) , 2015
2015
-
[26]
Sparse R-CNN: End-to-end object detection with learnable proposals,
P. Sun, R. Zhang, Y . Jiang, T. Kong, C. Xu, W. Zhan, M. Tomizuka, L. Li, Z. Yuan, C. Wang, and P. Luo, “Sparse R-CNN: End-to-end object detection with learnable proposals,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
-
[27]
End-to-end object detection with transformers,
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European Conference on Computer Vision (ECCV) , 2020
2020
-
[28]
Deformable DETR: Deformable transformers for end-to-end object detection,
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable DETR: Deformable transformers for end-to-end object detection,” in International Conference on Learning Representations (ICLR) , 2020
2020
-
[29]
Srp-uod: Multi-branch hybrid network framework based on structural re-parameterization for underwater small object detection,
J. Shi and W. Wu, “Srp-uod: Multi-branch hybrid network framework based on structural re-parameterization for underwater small object detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 2715–2719
2024
-
[30]
Sod-uav: Small object detection for unmanned aerial vehicle images via improved yolov7,
Y . Li, Y . Wang, Z. Ma, X. Wang, and Y . Tang, “Sod-uav: Small object detection for unmanned aerial vehicle images via improved yolov7,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024
2024
-
[31]
Small object detection on the water surface based on radar and camera fusion,
J. Zhu, Y . Yang, and Y . Cheng, “Small object detection on the water surface based on radar and camera fusion,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024
2024
-
[32]
Dynamic local and global context exploration for small object detection,
Z. Zhang, P. Gong, H. Sun, P. Wu, and X. Yang, “Dynamic local and global context exploration for small object detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023
2023
-
[33]
Parameter efficient transfer learning for various speech processing tasks,
S. Otake, R. Kawakami, and N. Inoue, “Parameter efficient transfer learning for various speech processing tasks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023
2023
-
[34]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Annual Conference on Neural Information Processing Systems (NeurIPS) , 2017, pp. 6000–6010
2017
-
[35]
Layer normalization,
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” in NeurIPS Deep Learning Symposium , 2016
2016
-
[36]
GLU variants improve transformer models,
N. Shazeer, “GLU variants improve transformer models,” arXiv preprint arXiv:2002.05202, 2020
2002 arXiv
-
[37]
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) , 2019
2019
-
[38]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in IEEE/CVF International Conference on Computer Vision (ICCV), 2021
2021
-
[39]
Objects365: A large-scale, high-quality dataset for object detection,
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun, “Objects365: A large-scale, high-quality dataset for object detection,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2019
2019
-
[40]
Mdetr-modulated detection for end-to-end multi- modal understanding,
A. Kamath et al., “Mdetr-modulated detection for end-to-end multi- modal understanding,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 1780–1790
2021
-
[41]
Kosmos-2: Grounding multimodal large language models to the world,
Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, and F. Wei, “Kosmos-2: Grounding multimodal large language models to the world,” arXiv preprint arXiv:2306.14824 , 2023
2023 arXiv
-
[42]
V3det: Vast vocabulary visual detection dataset,
J. Wang, P. Zhang, T. Chu, Y . Cao, Y . Zhou, T. Wu, B. Wang, C. He, and D. Lin, “V3det: Vast vocabulary visual detection dataset,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2023, pp. 19844– 19854
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.