Pith. sign in

REVIEW 2 cited by

Cross-domain Multi-modal Few-shot Object Detection via Rich Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.16188 v2 pith:3B42CWN4 submitted 2024-03-24 cs.CV

classification cs.CV
keywords featuremulti-modaltextdetectionfew-shotobjectrichcross-domain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cross-modal feature extraction and integration have led to steady performance improvements in few-shot learning tasks due to generating richer features. However, existing multi-modal object detection (MM-OD) methods degrade when facing significant domain-shift and are sample insufficient. We hypothesize that rich text information could more effectively help the model to build a knowledge relationship between the vision instance and its language description and can help mitigate domain shift. Specifically, we study the Cross-Domain few-shot generalization of MM-OD (CDMM-FSOD) and propose a meta-learning based multi-modal few-shot object detection method that utilizes rich text semantic information as an auxiliary modality to achieve domain adaptation in the context of FSOD. Our proposed network contains (i) a multi-modal feature aggregation module that aligns the vision and language support feature embeddings and (ii) a rich text semantic rectify module that utilizes bidirectional text feature generation to reinforce multi-modal feature alignment and thus to enhance the model's language understanding capability. We evaluate our model on common standard cross-domain object detection datasets and demonstrate that our approach considerably outperforms existing FSOD methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leveraging Language Prior for Infrared Small Target Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Language priors from GPT-4V, converted by CLIP into text embeddings, improve infrared small target detection when used only during training.

  2. Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A progressive curriculum that trains open-vocabulary detectors on low-ambiguity, high-signal cross-modal alignments first improves robustness to visual domain shifts, with modest, test-tuned gains.

Pith tools