Pith. sign in

REVIEW 6 cited by

C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14119 v3 pith:N23FZPJ3 submitted 2024-03-21 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords test-timecalibrationprompttuningc-tptclipdatadispersion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In deep learning, test-time adaptation has gained attention as a method for model fine-tuning without the need for labeled data. A prime exemplification is the recently proposed test-time prompt tuning for large-scale vision-language models such as CLIP. Unfortunately, these prompts have been mainly developed to improve accuracy, overlooking the importance of calibration, which is a crucial aspect for quantifying prediction uncertainty. However, traditional calibration methods rely on substantial amounts of labeled data, making them impractical for test-time scenarios. To this end, this paper explores calibration during test-time prompt tuning by leveraging the inherent properties of CLIP. Through a series of observations, we find that the prompt choice significantly affects the calibration in CLIP, where the prompts leading to higher text feature dispersion result in better-calibrated predictions. Introducing the Average Text Feature Dispersion (ATFD), we establish its relationship with calibration error and present a novel method, Calibrated Test-time Prompt Tuning (C-TPT), for optimizing prompts during test-time with enhanced calibration. Through extensive experiments on different CLIP architectures and datasets, we show that C-TPT can effectively improve the calibration of test-time prompt tuning without needing labeled data. The code is publicly accessible at https://github.com/hee-suk-yoon/C-TPT.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Training a policy to prefer responses by optimizing only its own low-confidence (high-surprisal) tokens improves alignment over uniform token optimization in SimPO and DPO.

  2. Measurement Plasticity: Sensor-Level Adaptation for Vision-Language Models

    cs.CV 2025-12 conditional novelty 5.0 of 10

    Selecting physical camera settings by feature-space affinity to the source domain and voting over multiple exposures outperforms digital-only test-time adaptation for vision-language models on sensor-shifted benchmarks.

  3. Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language Models

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    The submitted full text does not match the abstract, so the manuscript cannot be assessed as a coherent preprint.

  4. Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An online EM algorithm fits class-conditional Gaussians to the test stream from CLIP text-embedding initializations, improving test-time adaptation accuracy over prior methods on 15 benchmarks.

  5. Generalizing vision-language models to novel domains: A comprehensive survey

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A survey of VLM generalization literature organized by transferred module, with benchmark tables and a review of multimodal LLMs.

  6. CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text

    cs.CV 2025-08 reject novelty 2.0 of 10

    CLIPTime adds a classification head and a transformer-style regression head to CLIP embeddings, hitting 98.7% accuracy on synthetic fungi but with weak timestamp predictions, especially for spores.

Pith tools