Pith. sign in

REVIEW 2 cited by

Few-shot Action Recognition with Captioning Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10125 v1 pith:YGG3NHMD submitted 2023-10-16 cs.CV

classification cs.CV
keywords few-shotfoundationmodelscapfsarknowledgemultimodaltextaction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transferring vision-language knowledge from pretrained multimodal foundation models to various downstream tasks is a promising direction. However, most current few-shot action recognition methods are still limited to a single visual modality input due to the high cost of annotating additional textual descriptions. In this paper, we develop an effective plug-and-play framework called CapFSAR to exploit the knowledge of multimodal models without manually annotating text. To be specific, we first utilize a captioning foundation model (i.e., BLIP) to extract visual features and automatically generate associated captions for input videos. Then, we apply a text encoder to the synthetic captions to obtain representative text embeddings. Finally, a visual-text aggregation module based on Transformer is further designed to incorporate cross-modal spatio-temporal complementary information for reliable few-shot matching. In this way, CapFSAR can benefit from powerful multimodal knowledge of pretrained foundation models, yielding more comprehensive classification in the low-shot regime. Extensive experiments on multiple standard few-shot benchmarks demonstrate that the proposed CapFSAR performs favorably against existing methods and achieves state-of-the-art performance. The code will be made publicly available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LGA decomposes action labels and videos into three aligned atomic phases, then fuses text and video features to set a new state of the art in few-shot action recognition.

  2. How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction?

    cs.CV 2025-06 conditional novelty 6.0 of 10

    On a new NUGGET gesture dataset, skeleton-based HD-GCN beats vision foundation model V-JEPA (94.4% vs 90.1% top-1), while zero-shot Gemini Flash 2.0 achieves only 42.1%.

Pith tools