Pith. sign in

REVIEW 2 cited by

Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00860 v1 pith:Z7W4JK6H submitted 2024-04-01 cs.LG cs.CV

Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text Guidance

classification cs.LG cs.CV
keywords fine-tuningmodelsrobustlipsum-ftmodelzero-shotdatadistribution
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large-scale contrastive vision-language pre-trained models provide the zero-shot model achieving competitive performance across a range of image classification tasks without requiring training on downstream data. Recent works have confirmed that while additional fine-tuning of the zero-shot model on the reference data results in enhanced downstream performance, it compromises the model's robustness against distribution shifts. Our investigation begins by examining the conditions required to achieve the goals of robust fine-tuning, employing descriptions based on feature distortion theory and joint energy-based models. Subsequently, we propose a novel robust fine-tuning algorithm, Lipsum-FT, that effectively utilizes the language modeling aspect of the vision-language pre-trained models. Extensive experiments conducted on distribution shift scenarios in DomainNet and ImageNet confirm the superiority of our proposed Lipsum-FT approach over existing robust fine-tuning methods.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Model soups need only one ingredient

    cs.LG 2026-02 conditional novelty 6.0

    A single checkpoint, edited by splitting each layer's update with SVD and reweighting the high- and low-energy parts, reaches soup-level OOD robustness without multi-model training.

  2. Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs

    cs.CV 2026-07 reject novelty 4.0

    A simple domain-generalization training recipe (balanced real/tampered batches, late injection of a companion VLM domain, low learning rate) yields large pixel-level tampering-localization gains on out-of-distribution...