Pith. sign in

REVIEW 1 cited by

Enabling Calibration In The Zero-Shot Inference of Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.12748 v4 pith:XERCRQQC submitted 2023-03-11 cs.CV cs.LG

classification cs.CVcs.LG
keywords inferencecalibrationclipmodelszero-shotdatasetacrossarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Calibration of deep learning models is crucial to their trustworthiness and safe usage, and as such, has been extensively studied in supervised classification models, with methods crafted to decrease miscalibration. However, there has yet to be a comprehensive study of the calibration of vision-language models that are used for zero-shot inference, like CLIP. We measure calibration across relevant variables like prompt, dataset, and architecture, and find that zero-shot inference with CLIP is miscalibrated. Furthermore, we propose a modified version of temperature scaling that is aligned with the common use cases of CLIP as a zero-shot inference model, and show that a single learned temperature generalizes for each specific CLIP model (defined by a chosen pre-training dataset and architecture) across inference dataset and prompt choice.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prompting without Panic: Attribute-aware, Zero-shot, Test-Time Calibration

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Attribute-aware test-time prompt tuning with intra/inter-class text dispersion losses reduces average ECE from 11.7 to 4.11 across 11 fine-grained CLIP benchmarks.

Pith tools