Pith. sign in

REVIEW 3 cited by

Noise is an Efficient Learner for Zero-Shot Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.06019 v1 pith:45NYFH6H submitted 2025-02-09 cs.CV

classification cs.CV
keywords noisemodelstest-timetuningvisualadaptationadaptiveapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, test-time adaptation has garnered attention as a method for tuning models without labeled data. The conventional modus operandi for adapting pre-trained vision-language models (VLMs) during test-time primarily focuses on tuning learnable prompts; however, this approach overlooks potential distribution shifts in the visual representations themselves. In this work, we address this limitation by introducing Test-Time Noise Tuning (TNT), a novel method for handling unpredictable shifts in the visual space. TNT leverages, for the first time, a noise adaptation strategy that optimizes learnable noise directly in the visual input space, enabling adaptive feature learning from a single test sample. We further introduce a novel approach for inter-view representation alignment by explicitly enforcing coherence in embedding distances, ensuring consistent feature representations across views. Combined with scaled logits and confident view selection at inference, TNT substantially enhances VLM generalization and calibration, achieving average gains of +7.38% on natural distributions benchmark and +0.80% on cross-dataset evaluations over zero-shot CLIP. These improvements lay a strong foundation for adaptive out-of-distribution handling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adapting Vision-Language Models Without Labels: A Comprehensive Survey

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A survey that organizes unsupervised vision-language model adaptation by unlabeled-data availability into four paradigms: data-free transfer, domain transfer, episodic test-time, and online test-time adaptation.

  2. Decoupling Clinical and Class-Agnostic Features for Reliable Few-Shot Adaptation under Shift

    cs.CV 2025-09 reject novelty 4.0 of 10

    DRiFt explicitly decouples clinical from class-agnostic features in medical vision-language models and reports improved few-shot accuracy, but robustness under domain shift is not consistently supported.

  3. On the Robustness of Medical Vision-Language Models: Are they Truly Generalizable?

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Medical vision-language models lose accuracy on corrupted images; RobustMedCLIP, a few-shot LoRA-tuned BioMedCLIP, partially restores robustness on the new MediMeta-C benchmark.

Pith tools