Pith. sign in

REVIEW 2 cited by

Robust Domain Generalization for Multi-modal Object Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.05831 v1 pith:4YES4GWN submitted 2024-08-11 cs.CV cs.AI

classification cs.CVcs.AI
keywords acrossdomaingeneralizationlossrecognitionapproachesbackbonesclass-aware
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In multi-label classification, machine learning encounters the challenge of domain generalization when handling tasks with distributions differing from the training data. Existing approaches primarily focus on vision object recognition and neglect the integration of natural language. Recent advancements in vision-language pre-training leverage supervision from extensive visual-language pairs, enabling learning across diverse domains and enhancing recognition in multi-modal scenarios. However, these approaches face limitations in loss function utilization, generality across backbones, and class-aware visual fusion. This paper proposes solutions to these limitations by inferring the actual loss, broadening evaluations to larger vision-language backbones, and introducing Mixup-CLIPood, which incorporates a novel mix-up loss for enhanced class-aware visual fusion. Our method demonstrates superior performance in domain generalization across multiple datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Label-Decoupled Style Augmentation for Domain Generalization in Multi-Label Remote Sensing Scene Classification

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    Label-decoupled style augmentation raises multi-label remote-sensing domain-generalization mAP to 71.5%, beating ERM by 5.0 points and global style baselines by 1.3.

  2. Detecting and Classifying Defective Products in Images Using YOLO

    cs.CV 2024-12 reject novelty 2.0 of 10

    An unverifiable report that a YOLO variant with ResC2Net, SPPF, and PConv modules detects machine-part defects at mAP 0.91 without comparing to any baseline.

Pith tools