REVIEW 5 cited by
DynaPrompt: Dynamic Test-Time Prompt Tuning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Test-time prompt tuning enhances zero-shot generalization of vision-language models but tends to ignore the relatedness among test samples during inference. Online test-time prompt tuning provides a simple way to leverage the information in previous test samples, albeit with the risk of prompt collapse due to error accumulation. To enhance test-time prompt tuning, we propose DynaPrompt, short for dynamic test-time prompt tuning, exploiting relevant data distribution information while reducing error accumulation. Built on an online prompt buffer, DynaPrompt adaptively selects and optimizes the relevant prompts for each test sample during tuning. Specifically, we introduce a dynamic prompt selection strategy based on two metrics: prediction entropy and probability difference. For unseen test data information, we develop dynamic prompt appending, which allows the buffer to append new prompts and delete the inactive ones. By doing so, the prompts are optimized to exploit beneficial information on specific test data, while alleviating error accumulation. Experiments on fourteen datasets demonstrate the effectiveness of dynamic test-time prompt tuning.
Forward citations
Cited by 5 Pith papers
-
Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models
ZAEC anchors calibration to each sample's zero-shot entropy and selectively softens over-sharpened TTA predictions, reaching the lowest macro-average calibration error among evaluated post-hoc methods on ViT-B/16.
-
Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking
Adaptive Classifier-Free Guidance (A-CFG) re-masks low-confidence tokens in the unconditional input at each generation step, improving reasoning and planning accuracy for masked diffusion language models.
-
Progressive Scaling Visual Object Tracking
A progressive scaling training strategy with small-teacher distillation and masked-input alignment improves tracking accuracy and powers a new 12-dataset benchmark.
-
Measurement Plasticity: Sensor-Level Adaptation for Vision-Language Models
Selecting physical camera settings by feature-space affinity to the source domain and voting over multiple exposures outperforms digital-only test-time adaptation for vision-language models on sensor-shifted benchmarks.
-
Generalizing vision-language models to novel domains: A comprehensive survey
A survey of VLM generalization literature organized by transferred module, with benchmark tables and a review of multimodal LLMs.
Discussion (0). Continue with ORCID to comment.