Pith. sign in

REVIEW 4 major objections 5 minor 109 references

The paper claims that prompt-based adaptation of frozen vision transformers fails from unregulated layer-wise information flow, and that a layer-wise Information-Bottleneck regularizer — compression plus sufficiency plus a cross-layer path

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 06:09 UTC pith:ITZB5IXM

load-bearing objection Solid method contribution, but the headline VTAB-1k number is computed under a different aggregation rule than the baselines; fix that and the paper deserves refereeing. the 4 major comments →

arxiv 2607.21973 v1 pith:ITZB5IXM submitted 2026-07-24 cs.CV

Rethinking Layer-Wise Information Allocation for Vision Foundation Model Adaptation

classification cs.CV
keywords Information Bottleneckprompt tuningparameter-efficient adaptationvision transformerfrozen backbonelayer-wise regularizationtransfer learningshortcut learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that the weakness of prompt-based adaptation of frozen vision transformers is not insufficient prompt capacity but unregulated layer-wise information flow. It proposes Prompted Information Bottlenecks (PIB), a regularizer that encourages each prompted layer to compress redundant or nuisance information while preserving task-relevant discriminative evidence, and to evolve smoothly across depth. Across 34 datasets, PIB improves transfer accuracy — 92.1% on FGVC, 93.01% on HTA, and 77.33% on VTAB-1k — with about 0.35% trainable parameters on average, while making prompt scaling smoother and reducing reliance on shortcuts. A sympathetic reader would care because the paper offers a principled explanation of when and why prompt tuning fails, plus a lightweight practical fix.

Core claim

The paper claims that prompt-based VFM adaptation should be seen as a layer-wise information allocation problem rather than a prompt-design problem. It introduces PIB, which adds to the task loss layer-wise regularization: a compression term that minimizes off-diagonal correlation of normalized hidden states (a surrogate for reducing I(H_l;X)), a sufficiency term that minimizes intra-to-inter class scatter ratio (a surrogate for increasing I(H_l;Y)), and a cross-layer path regularizer that penalizes abrupt deterioration of the compression-sufficiency trade-off. Under this objective, effective prompts retain local evidence early and progressively discard nuisance and redundant detail; the res

What carries the argument

The central machinery is the Prompted Information Bottleneck (PIB) objective, decomposed into two tractable surrogates: compression via the off-diagonal energy of the normalized hidden-state correlation matrix, and sufficiency via the within-class versus between-class scatter ratio. A cross-layer path regularizer on layer quality scores enforces a coherent minimal-sufficient trajectory across depth, and optional learned per-layer routing gates rebalance compression versus sufficiency. The losses are computed from hidden states of the frozen transformer during training and are not used at inference.

Load-bearing premise

Everything rests on the assumption that the two surrogate losses — off-diagonal correlation for compression and class-scatter ratio for sufficiency — actually track what an information bottleneck would compress and preserve; the paper itself labels this a heuristic in its appendix.

What would settle it

Take PIB's two surrogates to a task where nuisance variability is not captured by hidden-state correlation and where task-relevant information does not show up as between-class scatter — for example dense prediction or image retrieval. If PIB then fails to beat vanilla visual prompt tuning, or hurts it, the assumption that the surrogates stand in for the information bottleneck is falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If PIB's view is right, prompt-tuning gains should concentrate on tasks where information misallocation accumulates most — the Structured group of VTAB-1k — and the paper reports exactly that: 63.18 versus 54.98 for VPT-Deep.
  • Non-monotonic behavior of prompt depth and length scaling is a symptom of unregulated information flow; PIB-regularized models scale more smoothly along both axes.
  • Reducing shortcut reliance: PIB improves foreground-only accuracy while degrading background-only accuracy, and lowers the average drop under corruption, domain shift, and few-shot transfer from 8.0 to 5.3 points.
  • These gains come with only about 0.35% trainable parameters on average, so the improvement is attributed to information regularization rather than added capacity.
  • The principle transfers across ViT-B/16, MAE- and MoCo-v3-pretrained ViT-B, and Swin-Base, indicating the layer-wise information-allocation framing is not tied to one architecture or pretraining recipe.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • An implication the paper leaves implicit: the same layer-wise compression-sufficiency framing could diagnose and regularize other parameter-efficient methods such as adapters or LoRA, since any adaptation that reshapes intermediate representations of a frozen model could be guided by a similar path constraint.
  • A testable extension: replace the correlation-matrix compression surrogate with a differentiable estimator of mutual information (for example, a variational bound) to see whether gains grow as the surrogate approaches the true IB objective; the paper's appendix concedes the current surrogates are heuristics.
  • The robustness results hint that PIB-like regularization could serve as a general shortcut-mitigation tool in fine-grained or distribution-shifted settings, independent of the prompt framework.
  • Because the path regularizer assumes smooth layer-wise improvement, tasks that genuinely benefit from abrupt representational transitions — for example some dense prediction or adversarial settings — may not fit; the paper acknowledges this limitation in its discussion.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes Prompted Information Bottlenecks (PIB), a regularization framework for frozen-backbone prompt tuning. PIB adds per-layer compression and sufficiency losses—implemented as off-diagonal correlation energy and inter/intra-class scatter ratio—plus a cross-layer path regularizer, and optionally a layer-routing gate. The authors claim state-of-the-art results on FGVC, HTA, and VTAB-1k (92.1%, 93.01%, 77.33%) with ~0.35% trainable parameters, smoother prompt scaling, reduced shortcut reliance, and more interpretable layer-wise information flow. The paper includes extensive experiments across multiple backbones, ablations, diagnostics, and a code link.

Significance. If the results hold, PIB would be a useful, lightweight contribution: it gives a principled information-allocation interpretation of prompt tuning and a practical regularizer that appears to improve transfer and robustness on a broad benchmark suite. The effort to connect layer-wise behavior to downstream robustness, and the availability of code, are clear strengths. However, the headline VTAB-1k number rests on an inconsistent aggregation rule, and the diagnostic evidence is partly circular because the same quantities are used as training objectives. The robustness protocol is also underspecified. These issues must be resolved before the empirical claims can be accepted.

major comments (4)
  1. [Table 1 / Appendix C.3 / Table S3] The reported VTAB-1k Mean of 77.33 for PIB is not the average over all 19 tasks as defined in Appendix C.3. It is the unweighted mean of the three group means: (82.98 + 85.83 + 63.18)/3 = 77.33. Using the per-task values in Table S3, the all-task mean is (7×82.98 + 4×85.83 + 8×63.18)/19 ≈ 75.24. Baselines such as VPT-D (69.43) and E2VPT (71.42) are computed as all-task means. The table therefore mixes aggregation protocols, and PIB's headline gain over VPT-D is overstated by roughly 2 points (7.9 vs 5.8). All rows must be recomputed under a single rule and the abstract/tables updated.
  2. [Eq. (6)/(9) vs. Eq. (17)/(20); §4.6, Appendix E] The diagnostic metrics used to support the information-allocation explanation—redundancy and class separability—are the same quantities as the compression and sufficiency training losses. The 'redundancy' diagnostic in Appendix E.2 is literally the off-diagonal correlation energy in Eq. (6), and the 'separability' diagnostic is the inverse of the sufficiency loss in Eq. (9). Showing that PIB reduces redundancy and improves separability relative to VPT is therefore expected from the objective. This is not independent evidence for an IB-style mechanism. The paper should either provide held-out, estimator-based measures (e.g., a non-parametric MI estimate) or explicitly weaken the interpretive claim.
  3. [Table 6] The shortcut-reliance and robustness experiment is advertised as a central contribution, but the protocol is not described. 'Std.', 'FG-only', 'BG-only', 'Corr.', 'Shift', and 'Few-shot' are undefined: no source dataset, no construction of foreground/background splits, no corruption or shift benchmarks, and no few-shot training details are given. Without this information the reported drops (8.0 → 5.3) cannot be reproduced or compared against other methods. Add a full protocol description, including dataset names, split construction, hyperparameters, and evaluation procedure.
  4. [Tables 1–6, Appendix A.2] The main tables report single numbers without standard deviations or seed counts, although Appendix A.2 says multiple random seeds are used and averages are reported. Since several conclusions rely on small margins (e.g., 92.1 vs 91.4 on FGVC, 93.01 vs 92.20 on HTA), missing uncertainty estimates make it impossible to tell whether differences are significant. Please report mean±std and the number of seeds for all main results, or clearly state the selection procedure.
minor comments (5)
  1. [References/figures] Figure cross-references are inconsistent: §4.4 discusses 'Figure 4' for scaling behavior, while Appendix D.3 calls it 'Figure 3'. Please unify figure numbering.
  2. [Equation numbering] Several key equations in the main text (e.g., compression and sufficiency losses in §3.3) are unnumbered, but later sections refer to equations. Add equation numbers throughout.
  3. [Appendix A.7] The layer-wise coefficient schedules λ_l, γ_l, α_l, δ_l are described qualitatively as 'lightweight' and 'depth-aware', but no actual values or normalization formula are given. Provide exact schedules or a precise pointer to code, since these coefficients define the method.
  4. [Related Work / Appendix I] Several cited works in Related Work and Appendix I (e.g., dataset distillation, gait recognition, composed retrieval) are not integrated into the narrative and seem tangential. Trimming these would improve focus.
  5. [Code URL] The repository name 'MM-26-PIB' suggests a conference submission rather than a stable arXiv code release. Please verify that the URL is active and contains the exact version used for the reported results.

Circularity Check

3 steps flagged

PIB's layer-wise diagnostic evidence reuses its own training losses: redundancy, separability, and path scores are rescaled versions of the compression, sufficiency, and path objectives, making those observations partly by construction; benchmark gains remain independent.

specific steps
  1. self definitional [Appendix E.2 (Eq. 17) vs. §3.3 (Eq. 5); used in §4.6 and Figure 1]
    "We reuse the same redundancy-based surrogate introduced in the main method. ... We measure redundancy using the normalized off-diagonal energy R_l = 1/(d(d-1)) ||C_l - I_d||^2_F,off. ... We then use the off-diagonal energy of C_l as the compression loss: L^{(l)}_{comp} = ||C_l - I_d||^2_F,off."

    PIB is trained to minimize L_comp at every prompted layer, so the diagnostic R_l is, up to normalization, the same objective value. Reporting that PIB has lower redundancy than VPT is therefore not independent evidence that PIB 'regulates information allocation'; it is a restatement that PIB reduced its own compression surrogate. The paper itself says the surrogate is not an exact mutual-information estimator, so the diagnostic adds little beyond checking that optimization worked.

  2. self definitional [Appendix E.2 (Eq. 20) vs. §3.3 (Eq. 9); used in §4.6]
    "For diagnostic clarity, we report the separability score in the 'higher-is-better' form D_l = S_inter/(S_intra+eps), which is the inverse counterpart of the sufficiency loss used during training."

    The sufficiency loss is L_suff = S_intra/(S_inter+eps). The diagnostic D_l is the reciprocal of this same ratio, so higher separability is exactly lower sufficiency loss. Since PIB is trained to minimize L_suff, the observed 'stronger class separability' is a direct consequence of optimizing that objective. It does not independently confirm that the learned representation is more task-sufficient; it confirms only that the chosen surrogate moved in the trained direction.

  3. self definitional [Appendix E.2 (Eqs. 21-22) vs. §3.4 (Eqs. 10-11); used in §4.6]
    "In the main method, we define the layer quality score q_l = L^{(l)}_{comp} + alpha_l L^{(l)}_{suff}. ... For ease of visualization, we convert this quantity into a 'higher-is-better' path score P_l = 1/(q_l+eps)."

    The path regularizer is L_path = sum max(0, q_{l+1}-q_l-delta_l), so PIB is explicitly trained to keep q_l from increasing across layers. The diagnostic 'more ordered path' is simply the inverse of q_l, i.e., the very quantity the path regularizer penalizes. Reporting that PIB has a smoother path score is a rescaled statement that the path loss was minimized, not an independent characterization of cross-layer information evolution.

full rationale

The paper's benchmark comparisons, robustness experiments, attention visualizations, and non-monotonic scaling analyses are independent of the identified circularity. The score is not higher because the central performance claims (FGVC 92.1, HTA 93.01, VTAB-1k group results, parameter efficiency, robustness gains) are not derived from the diagnostic metrics. However, the paper's mechanistic claim that PIB yields 'lower redundancy, stronger separability, and a more ordered cross-layer path' is supported in Appendix E by metrics that are literally rescaled versions of the losses PIB was trained to minimize. That is a concrete reduction: R_l, D_l, and P_l are the same quantities as L_comp, 1/L_suff, and 1/q_l, respectively. The paper does acknowledge in Appendix G and J that the surrogates are not exact information-theoretic quantities and that the path term is a structural prior, which mitigates the interpretive claim but does not remove the circularity of using the optimized objectives as post-hoc diagnostics. There is no load-bearing self-citation or imported uniqueness theorem; self-citations to the authors' prior prompt-tuning work are contextual. Separately, the VTAB-1k mean inconsistency (77.33 computed as a group-mean average versus ~75.24 as the stated all-task average) is a correctness/protocol issue affecting the headline number, but it is not a circularity in the derivation chain.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 0 invented entities

PIB's IB motivation is operationalized through heuristic surrogates; the central accuracy claim depends on these surrogates and a set of hand-chosen layer-wise coefficients. No new physical or architectural entities are introduced.

free parameters (5)
  • λ_l (per-layer compression weights) = not reported
    Hand-set depth-aware schedule (Appendix A.7); weights of the compression loss at each layer.
  • γ_l (per-layer sufficiency weights) = not reported
    Hand-set depth-aware schedule; weights of the sufficiency loss at each layer.
  • α_l (path-score mixing) = not reported
    Set to keep path score numerically balanced across layers; exact values absent.
  • δ_l (path tolerance margins) = not reported
    Chosen as a small non-negative tolerance; values absent.
  • η (path regularizer strength) = not reported
    Strength of L_path in the final objective; no value reported.
axioms (4)
  • domain assumption Reducing off-diagonal correlation of normalized hidden states approximates compression of task-irrelevant information.
    Section 3.3 uses it as the compression surrogate; no formal link to I(H_l;X).
  • domain assumption Class separability (intra/inter scatter ratio) approximates task sufficiency.
    Section 3.3 Eq. 9; acknowledged in Appendix G as tied to classification-style supervision.
  • domain assumption A smooth, progressively improving layer-quality sequence is a desirable prior for prompt tuning.
    Section 3.4 path regularizer; encoded as a structural prior, not derived.
  • ad hoc to paper Fixed depth-aware coefficient schedules transfer across backbones and datasets after normalizing loss magnitudes.
    Appendix A.7; no sensitivity analysis beyond the reported ablations.

pith-pipeline@v1.3.0-alltime-deepseek · 39013 in / 12237 out tokens · 107163 ms · 2026-08-01T06:09:10.274301+00:00 · methodology

0 comments
read the original abstract

Vision foundation models are increasingly reused as frozen backbones for downstream visual recognition, making parameter-efficient adaptation a central problem. Prompt-based adaptation, including Visual Prompt Tuning (VPT), provides a lightweight way to specialize these models, but its layer-wise behavior remains poorly understood: performance is sensitive to prompt depth, placement, and task distribution, and gains on standard in-domain benchmarks do not always translate into robust generalization. We argue that this limitation is not solely an optimization issue, but a layer-wise information allocation issue: existing prompt-based methods lack principled control over what prompt-conditioned representations should preserve, suppress, and propagate across depth. Inspired by the Information Bottleneck principle, we introduce Prompted Information Bottlenecks (PIB), a framework that regularizes layer-wise compression-sufficiency trade-offs and promotes a more coherent cross-layer information path. The key idea is that effective adaptation should be minimal yet sufficient, retaining task-relevant local evidence in earlier layers while progressively discarding nuisance factors and redundant details in deeper layers. Extensive experiments show that PIB achieves strong performance across 34 datasets, reaching 92.1% on FGVC, 93.01% on HTA, and 77.33% on VTAB-1k, while tuning only 0.35% parameters on average across the main settings. Beyond benchmark accuracy, PIB helps explain the non-monotonic behavior of prompt capacity scaling, reduces shortcut reliance, and improves robustness under distribution shift and fine-grained recognition settings. These results position PIB as both a practical method and an information-allocation perspective for adapting frozen vision foundation models. Our code is available at https://github.com/itsnotacie/MM-26-PIB

Figures

Figures reproduced from arXiv: 2607.21973 by Aiden Zhao, Hao Xu, Lin Zhao, Tianyang Wang, Xi Xiao, Yingli Tian, Yu Li, Yunbei Zhang, Yuqi Li.

Figure 1
Figure 1. Figure 1: A motivating diagnosis of layer-wise prompt behav [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of our proposed Prompted Information Bottleneck (PIB). A: prompt-based adaptation inserts trainable [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Average attention logits across transformer depth. We compare the layer-wise attention patterns of VPT and PIB [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Depth and length scaling behavior of prompt-based [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative visualization of prompt attention. Compared with VPT and ViaPT, our method produces more concentrated [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

109 extracted references · 3 canonical work pages

  1. [1]

    Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy. 2016. Deep variational information bottleneck.arXiv preprint arXiv:1612.00410(2016). 3

  2. [2]

    Reza Azad, Abdur R Fayjie, Claude Kauffmann, Ismail Ben Ayed, Marco Pedersoli, and Jose Dolz. 2021. On the texture bias for few-shot cnn segmentation. In Proceedings of the IEEE/CVF winter conference on applications of computer vision. 2674–2683. 3

  3. [3]

    Han Cai, Chuang Gan, Ligeng Zhu, and Song Han. 2020. Tinytl: Reduce memory, not parameters for efficient on-device learning.Advances in Neural Information Processing Systems33 (2020), 11285–11297. 2, 6

  4. [4]

    Gal Chechik, Amir Globerson, Naftali Tishby, and Yair Weiss. 2003. Information bottleneck for Gaussian variables.Advances in Neural Information Processing Systems16 (2003). 3

  5. [5]

    Kai Chen, Han Yu, Ye Wang, Xin Song, Xiaojuan Zhao, Yalong Xie, Liqun Gao, and Aiping Li. 2025. Temporal knowledge graph extrapolation with subgraph information bottleneck.Expert Systems with Applications268 (2025), 126226. 3

  6. [6]

    Shoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang, Yibing Song, Jue Wang, and Ping Luo. 2022. Adaptformer: Adapting vision transformers for scalable visual recognition.Advances in Neural Information Processing Systems35 (2022), 16664–16678. 2, 6, 16, 27

  7. [7]

    Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. 2020. Improved baselines with momentum contrastive learning.arXiv preprint arXiv:2003.04297(2020). 2, 6

  8. [8]

    Xinlei Chen, Saining Xie, and Kaiming He. 2021. An empirical study of training self-supervised vision transformers. InProceedings of the IEEE/CVF international conference on computer vision. 9640–9649. 6

  9. [9]

    Zhiwei Chen, Yupeng Hu, Zhiheng Fu, Zixu Li, Jiale Huang, Qinlei Huang, and Yinwei Wei. 2026. INTENT: Invariance and Discrimination-aware Noise Mitiga- tion for Robust Composed Image Retrieval. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 20463–20471. doi:10.1609/aaai.v40i25.39181 28

  10. [10]

    Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, Guozhi Qiu, Weili Guan, and Liqiang Nie. 2026. EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge.arXiv preprint arXiv:2605.24500(2026). 29

  11. [11]

    Hyunhee Chung and Kyung Ho Park. 2022. Shape prior is not all you need: discovering balance between texture and shape bias in CNN. InProceedings of the Asian Conference on Computer Vision. 4160–4175. 3

  12. [12]

    Rajshekhar Das, Yonatan Dukler, Avinash Ravichandran, and Ashwin Swami- nathan. 2023. Learning expressive prompting with residuals for vision transform- ers. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3366–3377. 2, 6

  13. [13]

    Shaohua Dong, Yunhe Feng, Qing Yang, Yan Huang, Dongfang Liu, and Heng Fan. 2024. Efficient multimodal semantic segmentation via dual-prompt learning. In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 14196–14203. 2, 6

  14. [14]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. [n. d.]. An Image is Worth 16x16 Words: Trans- formers for Image Recognition at Scale. InInternational Conference on Learning Representations. 1, 2, 6

  15. [15]

    Gal Elidan, Nir Friedman, and David Maxwell Chickering. 2005. Learning Hidden Variable Networks: The Information Bottleneck Approach.Journal of Machine Learning Research6, 1 (2005). 3

  16. [16]

    Marco Federici, Anjan Dutta, Patrick Forré, Nate Kushman, and Zeynep Akata

  17. [17]

    Ziv Goldfeld and Yury Polyanskiy. 2020. The information bottleneck problem and its applications in machine learning.IEEE Journal on Selected Areas in Information Theory1, 1 (2020), 19–38. 3

  18. [18]

    Cheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao, Wenguan Wang, Siyuan Qi, and Dongfang Liu. 2023. E 2 VPT: An Effective and Efficient Approach for Visual Prompt Tuning. In2023 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, 17445–17456. 2, 6, 7, 16, 27

  19. [19]

    Cheng Han, Qifan Wang, Yiming Cui, Wenguan Wang, Lifu Huang, Siyuan Qi, and Dongfang Liu. [n. d.]. Facing the Elephant in the Room: Visual Prompt Tuning or Full finetuning?. InThe Twelfth International Conference on Learning Representations. 2

  20. [20]

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick

  21. [21]

    Katherine Hermann, Ting Chen, and Simon Kornblith. 2020. The origins and prevalence of texture bias in convolutional neural networks.Advances in neural information processing systems33 (2020), 19000–19015. 3

  22. [22]

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. InInternational conference on machine learning. PMLR, 2790–2799. 16

  23. [23]

    Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. [n. d.]. LoRA: Low-Rank Adaptation of Large Language Models. InInternational Conference on Learning Representations. 2, 6, 16, 27

  24. [24]

    Qidong Huang, Xiaoyi Dong, Dongdong Chen, Weiming Zhang, Feifei Wang, Gang Hua, and Nenghai Yu. 2023. Diversity-aware meta visual prompting. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10878–10887. 2, 6, 16

  25. [25]

    Eugenia Iofinova, Alexandra Peste, Mark Kurtz, and Dan Alistarh. 2022. How well do sparse imagenet models transfer?. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12266–12276. 6, 7

  26. [26]

    Akshay Vivek Jagadeesh and Margaret Livingstone. 2024. Texture bias in primate ventral visual cortex. InICLR 2024 Workshop on Representational Alignment. 3

  27. [27]

    Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. 2022. Visual prompt tuning. InEuro- pean conference on computer vision. Springer, 709–727. 1, 2, 6, 7, 16, 27

  28. [28]

    Can Jin, Ying Li, Mingyu Zhao, Shiyu Zhao, Zhenting Wang, Xiaoxiao He, Ligong Han, Tong Che, and Dimitris Metaxas. 2025. Lor-VP: Low-rank visual prompting for efficient vision model adaptation. InInternational Conference on Learning Representations, Vol. 2025. 73004–73021. 2, 6, 7

  29. [29]

    Yanshu Li, Jiaqian Li, Kuai Yu, Xi Xiao, Dongfang Liu, Tianyang Wang, and Ruixiang Tang. 2026. Personalize your large vision-language models with in- context prompt tuning.arXiv preprint arXiv:2605.31513(2026). 2

  30. [30]

    Yuqi Li, Siwei Meng, Chuanguang Yang, Weilun Feng, Junming Liu, Zhulin An, Yikai Wang, and Yingli Tian. 2026. A comprehensive survey of interaction techniques in 3d scene generation. (2026). 2

  31. [31]

    Zixu Li, Zhiwei Chen, Zhiheng Fu, Wenbo Wang, Yupeng Hu, Weili Guan, and Liqiang Nie. 2026. OmniEgo-R 2: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026.arXiv preprint arXiv:2605.24481 (2026). 29

  32. [32]

    Zhiheng Li, Ivan Evtimov, Albert Gordo, Caner Hazirbas, Tal Hassner, Cris- tian Canton Ferrer, Chenliang Xu, and Mark Ibrahim. 2023. A whac-a-mole dilemma: Shortcuts come in multiples where mitigating one amplifies others. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion. 20071–20082. 3

  33. [33]

    Zixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang, Guozhi Qiu, Zhiheng Fu, and Meng Liu. 2026. ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 23373–23381. doi:10.1609/aaai.v40i28. 39507 28

  34. [34]

    Zixu Li, Yupeng Hu, Zhiwei Chen, Mingyu Zhang, Zhiheng Fu, and Liqiang Nie

  35. [35]

    Zixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen, Yongqi Li, and Liqiang Nie. 2026. TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 24421–24442. doi:10.18653/ v1/2026.acl-long.1121 28

  36. [36]

    Li Liu, Jie Chen, Paul Fieguth, Guoying Zhao, Rama Chellappa, and Matti Pietikäi- nen. 2019. From BoW to CNN: Two decades of texture representation for texture classification.International Journal of Computer Vision127, 1 (2019), 74–109. 3

  37. [37]

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer us- ing shifted windows. InProceedings of the IEEE/CVF international conference on computer vision. 10012–10022. 7

  38. [38]

    Oren Nuriel, Sharon Fogel, and Ron Litman. 2022. Textadain: Paying attention to shortcut learning in text recognizers. InEuropean Conference on Computer Vision. Springer, 427–445. 3

  39. [39]

    Ziqi Pan, Li Niu, Jianfu Zhang, and Liqing Zhang. 2021. Disentangled information bottleneck. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 9285–9293. 3

  40. [40]

    Wenjie Pei, Tongqi Xia, Fanglin Chen, Jinsong Li, Jiandong Tian, and Guangming Lu. 2024. SA 2VP: Spatially Aligned-and-Adapted Visual Prompt. InProceedings of the AAAI conference on artificial intelligence, Vol. 38. 4450–4458. 2, 6, 7

  41. [41]

    Zoe Piran, Ravid Shwartz-Ziv, and Naftali Tishby. 2020. The dual information bottleneck.arXiv preprint arXiv:2006.04641(2020). 3

  42. [42]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learning. PmLR, 8748–8763. 1

  43. [43]

    Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017. Learning mul- tiple visual domains with residual adapters.Advances in neural information processing systems30 (2017). 2, 6, 7

  44. [44]

    Li Ren, Chen Chen, Liqiang Wang, and Kien Hua. 2025. DA-VPT: Semantic-guided visual prompt tuning for vision transformers. InProceedings of the Computer Vision and Pattern Recognition Conference. 4353–4363. 2, 6 Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Li et al

  45. [45]

    Andrew M Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D Tracey, and David D Cox. 2019. On the information bottleneck theory of deep learning.Journal of Statistical Mechanics: Theory and Experiment2019, 12 (2019), 124020. 3

  46. [46]

    2002.The information bottleneck: Theory and applications

    Noam Slonim. 2002.The information bottleneck: Theory and applications. Ph. D. Dissertation. Hebrew University of Jerusalem Jerusalem, Israel

  47. [47]

    Noam Slonim and Naftali Tishby. 1999. Agglomerative information bottleneck. Advances in neural information processing systems12 (1999). 3

  48. [48]

    Pirzada Suhail, Vrinda Goel, and Amit Sethi. 2026. Shortcut learning susceptibility in vision classifiers (student abstract). InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 41390–41392. 3

  49. [49]

    Yehui Tang, Kai Han, Chang Xu, An Xiao, Yiping Deng, Chao Xu, and Yunhe Wang. 2021. Augmented shortcuts for vision transformers.Advances in Neural Information Processing Systems34 (2021), 15316–15327

  50. [50]

    Johannes Theodoridis, Jessica Hofmann, Johannes Maucher, and Andreas Schilling. 2022. Trapped in texture bias? a large scale comparison of deep instance segmentation. InEuropean Conference on Computer Vision. Springer, 609–627. 3

  51. [51]

    Naftali Tishby, Fernando C Pereira, and William Bialek. 2000. The information bottleneck method.arXiv preprint physics/0004057(2000). 2, 3

  52. [52]

    Naftali Tishby and Noga Zaslavsky. 2015. Deep learning and the information bottleneck principle. In2015 ieee information theory workshop (itw). Ieee, 1–5. 2, 3

  53. [53]

    Matias Vera, Leonardo Rey Vega, and Pablo Piantanida. 2018. Collaborative information bottleneck.IEEE Transactions on Information Theory65, 2 (2018), 787–815. 3

  54. [54]

    Zhibin Wan, Changqing Zhang, Pengfei Zhu, and Qinghua Hu. 2021. Multi- view information-bottleneck representation learning. InProceedings of the AAAI conference on artificial intelligence, Vol. 35. 10085–10092. 3

  55. [55]

    Junze Wang, Lei Fan, Dezheng Zhang, Weipeng Jing, Donglin Di, Yang Song, Sidong Liu, and Cong Cong. 2026. Visual Prompt-Agnostic Evolution.arXiv preprint arXiv:2601.20232(2026). 2

  56. [56]

    Xin Wang, Nicolas D Georganas, and Emil M Petriu. 2010. Fabric texture analysis using computer vision techniques.IEEE transactions on instrumentation and measurement60, 1 (2010), 44–56. 3

  57. [57]

    Yuzhu Wang, Lechao Cheng, Chaowei Fang, Dingwen Zhang, Manni Duan, and Meng Wang. 2024. Revisiting the power of prompt for visual tuning. In Proceedings of the 41st International Conference on Machine Learning. 50233–50247. 2, 6

  58. [58]

    Tailin Wu, Ian Fischer, Isaac L Chuang, and Max Tegmark. 2020. Learnability for the information bottleneck. InUncertainty in Artificial Intelligence. PMLR, 1050–1060. 3

  59. [59]

    Tailin Wu, Hongyu Ren, Pan Li, and Jure Leskovec. 2020. Graph information bottleneck.Advances in Neural Information Processing Systems33 (2020), 20437– 20448. 3

  60. [60]

    ZhengXian Wu, Hangrui Xu, Kai Shi, Zhuohong Chen, Yunyao Yu, Chuanrui Zhang, Zirui Liao, Jun Yang, Zhenyu Yang, Haonan Lu, et al . 2026. ProMSA: Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering.arXiv preprint arXiv:2606.27974(2026). 28

  61. [61]

    Zhengxian Wu, Chuanrui Zhang, Shen’Ao Jiang, Hangrui Xu, Zirui Liao, Luyuan Zhang, Li Huaqiu, Peng Jiao, and Haoqian Wang. 2026. Language-guided and motion-aware gait representation for generalizable recognition. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 10871–10878. 28

  62. [62]

    Zhengxian Wu, Chuanrui Zhang, Hangrui Xu, Peng Jiao, and Haoqian Wang

  63. [63]

    Xi Xiao, Xingjian Li, Yunbei Zhang, Cheng Han, Tianming Liu, Tianyang Wang, Runmin Jiang, Jihun Hamm, Xiao Wang, and Min Xu. 2026. Layer-specific prompt fusion discovery via differentiable search in vision foundation models.arXiv preprint arXiv:2606.26379(2026). 2

  64. [64]

    Xi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu, Zhuxuanzi Wang, Yanshu Li, Lin Zhao, Guosheng Hu, Tianyang Wang, and Hao Xu. 2026. Not all directions matter: Towards structured and task-aware low-rank model adaptation. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2132–2154. 3

  65. [65]

    Xi Xiao, Wentao Wang, Jiacheng Xie, Lijing Zhu, Gaofei Chen, Zhengji Li, Tianyang Wang, and Min Xu. 2024. HGTDP-DTA: Hybrid graph-transformer with dynamic prompt for drug-target binding affinity prediction. InInternational Conference on Neural Information Processing. Springer, 340–354. 2

  66. [66]

    Xi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang, Xiao Wang, Yuxiang Wei, Jihun Hamm, and Min Xu. 2025. Visual instance-aware prompt tuning. InPro- ceedings of the 33rd ACM International Conference on Multimedia. 2880–2889. 2, 6, 7

  67. [67]

    Xi Xiao, Yunbei Zhang, Yanshuh Li, Xingjian Li, Tianyang Wang, Jihun Hamm, Xiao Wang, and Min Xu. 2025. Visual variational autoencoder prompt tuning. arXiv preprint arXiv:2503.17650(2025). 2

  68. [68]

    Xi Xiao, Yunbei Zhang, Lin Zhao, Yiyang Liu, Xiaoying Liao, Zheda Mai, Xingjian Li, Xiao Wang, Hao Xu, Jihun Hamm, et al. 2025. Prompt-based adaptation in large-scale vision models: A survey.arXiv preprint arXiv:2510.13219(2025). 2

  69. [69]

    Huangbiao Xu, Xiao Ke, Yuezhou Li, Rui Xu, Huanqi Wu, Xiaofeng Lin, and Wenzhong Guo. 2024. Vision-Language Action Knowledge Learning for Semantic- Aware Action Quality Assessment. InECCV. 2

  70. [70]

    Huangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu, Yuezhou Li, and Wenzhong Guo

  71. [71]

    Huangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu, Yuezhou Li, Peirong Xu, and Wen- zhong Guo. 2025. Dancefix: An exploration in group dance neatness assessment through fixing abnormal challenges of human pose. InAAAI. 2

  72. [72]

    Huangbiao Xu, Huanqi Wu, Xiao Ke, Yuezhou Li, Rui Xu, and Wenzhong Guo

  73. [73]

    Hangrui Xu, Zhengxian Wu, Chuanrui Zhang, Zhuohong Chen, Zhifang Liu, Peng Jiao, and Haoqian Wang. 2026. Psgait: Gait recognition using parsing skeleton. InICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 10427–10431. 28

  74. [74]

    Jiawei Xu, Qiangqiang Zhou, Zhouping Li, Yanjiao Shi, Yugen Yi, and Jiacong Yu

  75. [75]

    Language-Guided Audio-Visual Learning for Long-Term Sports Assessment. InCVPR. 2

  76. [76]

    Jiawei Xu, Qiangqiang Zhou, Dandan Zhu, Yong Chen, Yugen Yi, and Xiaoqi Zhao. 2026. TP-Seg: Task-Prototype Framework for Unified Medical Lesion Segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5452–5462. 28

  77. [77]

    Zhipei Xu, Xuanyu Zhang, Runyi Li, Zecheng Tang, Qing Huang, and Jian Zhang

  78. [78]

    Quality-Guided Vision-Language Learning for Long-Term Action Quality Assessment.IEEE Transactions on Multimedia(2025). 2

  79. [79]

    Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. 2014. How transferable are features in deep neural networks?Advances in neural information processing systems27 (2014). 2, 6

  80. [80]

    Haonan Yuan, Qingyun Sun, Xingcheng Fu, Cheng Ji, and Jianxin Li. 2024. Dy- namic graph information bottleneck. InProceedings of the ACM Web Conference

Showing first 80 references.