Pith. sign in

Paper Citation Record · LEDGER

One Last Attention for Your Vision-Language Model

As of 22 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 0 inbound Pith citation observations for arXiv:2507.15480.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15480 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:42:16.515695Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 321dcd3b-9e0e-457d-a8cd-e4970c7b7d33 · outbound

This paper cites Align your prompts: Test-time prompting with distribution align- ment for zero-shot generalization.

One Last Attention for Your Vision-Language Model Align your prompts: Test-time prompting with distribution align- ment for zero-shot generalization

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.970643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:11.650459Z digest=sha256:e18a72b38e223be82d4152ce07fe1b2b51d618a301ddc20ba8d54f6e907da8de

Observation dbc2113f-d2af-4331-af92-f33e5587c5db · outbound

This paper cites Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.

One Last Attention for Your Vision-Language Model Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.964204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:11.726490Z digest=sha256:6bea36a1f2218ce47a6aaa2025e1e8992454fb3517aaea8c7edf158d1856702d

Observation 66259e06-ce9c-44cf-aea3-cf82d209198f · outbound

This paper cites Food-101–mining discriminative components with random forests.

One Last Attention for Your Vision-Language Model Food-101–mining discriminative components with random forests

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:11.878130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:11.878130Z digest=sha256:142baed18920231baa619810eb1950f338a2043beddbcd643108a7583ef59170

Observation ec0bcd0c-b9ed-43de-a3be-9facfa120c7a · outbound

This paper cites Improved test-time adaptation for domain gen- eralization.

One Last Attention for Your Vision-Language Model Improved test-time adaptation for domain gen- eralization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.953617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:12.052169Z digest=sha256:803056d13994f6aa998b6d27bf2e4a4e6e29bcf74de68b77f3ebbdd452eae370

Observation 8dd90c77-46eb-418a-b889-cce7ab4a1de9 · outbound

This paper cites Domain generalization via rationale invariance.

One Last Attention for Your Vision-Language Model Domain generalization via rationale invariance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.946889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:12.220184Z digest=sha256:6256f28240827b1746ef00627e613e47a24b7dbd4ec4277ab3652bb722203ce2

Observation af2f4e97-fc3f-4daf-beb4-f14c7cc3646b · outbound

This paper cites Lfme: A simple framework for learning from multiple experts in domain generalization.NeurIPS, 2024.

One Last Attention for Your Vision-Language Model Lfme: A simple framework for learning from multiple experts in domain generalization.NeurIPS, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.939899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:12.385901Z digest=sha256:a51fa8a24994de23e05453d8cdf7d5ac368245e18778ed752dc075ab3c7be433

Observation 65a4ae63-dc37-4d90-b6eb-455391ffeda7 · outbound

This paper cites A causal inspired early-branching structure for domain generalization.IJCV, 132(9):4052–4072, 2024.

One Last Attention for Your Vision-Language Model A causal inspired early-branching structure for domain generalization.IJCV, 132(9):4052–4072, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.933568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:12.550962Z digest=sha256:9c88d5cdb1ffadd6141167754a16aa0ebe75391a0a6b513c5e5b5c0ceedebdf3

Observation 6aab2e37-5c36-4390-b1a9-5944e6855af2 · outbound

This paper cites Vision transformer adapter for dense predictions.

One Last Attention for Your Vision-Language Model Vision transformer adapter for dense predictions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.927106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:12.677415Z digest=sha256:e4ef9e9527f946ce4965fa8fcb1b3ccd42e564eb9851f082e8094e88b9001b08

Observation 253661df-edc8-4635-bca2-256884d2695e · outbound

This paper cites Describing textures in the wild.

One Last Attention for Your Vision-Language Model Describing textures in the wild

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.920711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:12.812714Z digest=sha256:619c707c184c0fe26aa0cddb5ce60e364ae8785e2977fc5045c7d47527c474d4

Observation 1cc54bb4-8580-4eb3-8317-0bfbaed4e9aa · outbound

This paper cites John Wiley & Sons, 1999.

One Last Attention for Your Vision-Language Model John Wiley & Sons, 1999

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.914220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:12.992561Z digest=sha256:6aee9a3a38b6ea670a71fd5bbb7c17b20d322a0f800bb06d9e8ab6b02dbc57fc

Observation c8913d0f-2b25-4327-885d-1f98ac450c8e · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

One Last Attention for Your Vision-Language Model Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:13.113425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:13.113425Z digest=sha256:2f987feb53ab23e0fb9c065ff9bb2135600448a5637744fb40d798a14a95c1dc

Observation e19a8c1a-3472-4ffb-a77d-1483c2b7f8e2 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

One Last Attention for Your Vision-Language Model An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.904261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:13.269657Z digest=sha256:27545876bd1ce17b341c8dc41d18cf83c23dff5cbdf5a1235df30dfb9d1ba6bb

Observation b77fccd7-91ff-4a65-b4f8-0ace6adb402b · outbound

This paper cites Magma–multimodal augmentation of generative models through adapter-based finetuning.

One Last Attention for Your Vision-Language Model Magma–multimodal augmentation of generative models through adapter-based finetuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.897187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:13.459632Z digest=sha256:71bf373720e94769c1339802e97e1ab7c2a3d8409626c11b996c60ae00a80340

Observation 63c7dc0a-99c8-4bbd-8d14-a5b0aae12bf4 · outbound

This paper cites Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories.

One Last Attention for Your Vision-Language Model Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:13.614063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:13.614063Z digest=sha256:05e425aba5821aeb5fb2acd517abf46545168eaec86d14a661e5a8abc74d66f3

Observation 48dc7b24-8868-4700-93c9-f2e04161c3ad · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.IJCV, 132(2):581–595, 2024.

One Last Attention for Your Vision-Language Model Clip-adapter: Better vision-language models with feature adapters.IJCV, 132(2):581–595, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.887176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:13.758312Z digest=sha256:74a585ef96f93f20ddfc6a2d5d78d074305d8ea96c45f909473563d285857015

Observation bfeb9dd3-97c7-405c-9b5f-1a680497910f · outbound

This paper cites Finetune like you pretrain: Im- proved finetuning of zero-shot vision models.

One Last Attention for Your Vision-Language Model Finetune like you pretrain: Im- proved finetuning of zero-shot vision models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.879928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:13.910566Z digest=sha256:583fab9425bc52ff77776eb39ffdb9e64e17b26fb003ea9e0062a9d03297769c

Observation 30bef324-0875-4ee7-ba15-1c57cab3c218 · outbound

This paper cites Deep residual learning for image recognition.

One Last Attention for Your Vision-Language Model Deep residual learning for image recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:14.055195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:14.055195Z digest=sha256:0cf5e6cbccff55b4e018fd4580c973c946870379a06cedb14cdaed3f048dbe34

Observation e558ee89-7489-4ad1-a3ac-1ae16a029c01 · outbound

This paper cites an unresolved cited work.

One Last Attention for Your Vision-Language Model Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:14.224047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:14.224047Z digest=sha256:b235cb96dbf1c8236c86f0a718f075b123d7bba378b04ef6e4b9516206a3b4b2

Observation c9983d55-8e37-4b6c-b9e2-c95356ff1d3e · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.ICCV, 2021.

One Last Attention for Your Vision-Language Model The many faces of robustness: A critical analysis of out-of-distribution generalization.ICCV, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.865977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:14.361257Z digest=sha256:74198489659312ac6787e98a8be0b71e80d182fe62067b3843ba039d64498d44

Observation 988166e3-e1f1-4edf-97e3-3ad96b945bfc · outbound

This paper cites Natural adversarial examples.

One Last Attention for Your Vision-Language Model Natural adversarial examples

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.859038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:14.504720Z digest=sha256:cc9967be6d48d5d8b7d4748148602b355db18e51851f98a1e0763916dd1de98b

Observation 2786a37f-8214-4474-853d-a8637647470e · outbound

This paper cites Open- clip, 2021.

One Last Attention for Your Vision-Language Model Open- clip, 2021

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.852641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:14.637867Z digest=sha256:58d17c80e8068b10e836d66cc1b4a1f40d37c89e6b93157285d13401e995a0d0

Observation a6a73f54-886b-4776-86a7-b67ebc3b9008 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

One Last Attention for Your Vision-Language Model Scaling up visual and vision-language representation learning with noisy text supervision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.846580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:14.799608Z digest=sha256:f7b14d0102f9c4d74de959abfcec4728b541d50a21533ad41273f2db99480428

Observation 6bb61b29-3ceb-4454-bbde-9bf854be401b · outbound

This paper cites Vi- sual prompt tuning.

One Last Attention for Your Vision-Language Model Vi- sual prompt tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.840036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:14.965203Z digest=sha256:77466d90f8f044f350c16e673a83c6a2b9cf07ffc2de98146384212e5ed99795

Observation 15eb36b3-c394-4a95-bc76-f5b6466dfb38 · outbound

This paper cites Maple: Multi-modal prompt learning.

One Last Attention for Your Vision-Language Model Maple: Multi-modal prompt learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.833700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:15.153022Z digest=sha256:c684c13a5bb9d1917835b79307114667418e9edb4143a10bbbccb4a20e8b47ab

Observation ea2d5154-40d3-4ea3-8bbb-33e6ea943c4c · outbound

This paper cites 3d object representations for fine-grained categorization.

One Last Attention for Your Vision-Language Model 3d object representations for fine-grained categorization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:15.303596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:15.303596Z digest=sha256:01ca9399e638e178a3f37e237e4f0e2e552702205fc6a0b499af08217bca3a43

Observation 56fa0de1-2420-437c-8f1b-e48b6b17d613 · outbound

This paper cites Fine-tuning can distort pretrained fea- tures and underperform out-of-distribution.

One Last Attention for Your Vision-Language Model Fine-tuning can distort pretrained fea- tures and underperform out-of-distribution

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.823263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:15.455137Z digest=sha256:5cc9e51632549c2793bfab1805e82691865e7041c343400ae999f2dce104b07a

Observation e25a7b87-2e78-4a2e-bbd8-28c9b150114f · outbound

This paper cites Su- pervision exists everywhere: A data efficient contrastive language-image pre-training paradigm.

One Last Attention for Your Vision-Language Model Su- pervision exists everywhere: A data efficient contrastive language-image pre-training paradigm

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.817023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:15.619578Z digest=sha256:e4f80712880c813acf3861e0c6d5bfdfe7ef634ec78500dbc64463f12825fbac

Observation 00737bd9-52cf-4d5f-9bb6-4760ecf175a5 · outbound

This paper cites Ttt++: When does self-supervised test-time training fail or thrive? InNeurIPS, 2021.

One Last Attention for Your Vision-Language Model Ttt++: When does self-supervised test-time training fail or thrive? InNeurIPS, 2021

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.809627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:15.750028Z digest=sha256:0e571cc435077d57c9097157f01c0bb589a2ce08ea8cfac2b75ea5c78b94ab4e

Observation 053395b7-fb28-4de3-ba59-055f091220e1 · outbound

This paper cites Decoupled weight decay regularization.

One Last Attention for Your Vision-Language Model Decoupled weight decay regularization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.802900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:15.980425Z digest=sha256:a1bcb4fa3c8bffe7b7c0cb5be8aca9113069eadfff3d92af33463735cc70b98c

Observation a2ac80c5-e7ab-46a5-95d3-68808ae6df8e · outbound

This paper cites A Simple Long-Tailed Recognition Baseline via Vision-Language Model.

One Last Attention for Your Vision-Language Model A Simple Long-Tailed Recognition Baseline via Vision-Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.043292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.043292Z digest=sha256:0af0735871cde56d82830fc182326264e87df2049423fd183a3f763228088357

Observation 36181347-35d0-41b1-9d1a-f95f1aa1c8d8 · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

One Last Attention for Your Vision-Language Model Fine-Grained Visual Classification of Aircraft

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.121014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.121014Z digest=sha256:dfea561da31c466b3774fb5d6b168660775c351b21373a278f7a5a353b476e07

Observation 020b82f6-5802-4349-bdb6-8cce96cb92ea · outbound

This paper cites Automated flower classification over a large number of classes.

One Last Attention for Your Vision-Language Model Automated flower classification over a large number of classes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.244453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.244453Z digest=sha256:093f1e943ebf0c98af0eb8a9d865606ba01ddbb7a40f95d3106aa314d328bce1

Observation c6adfb22-443e-4f5c-93f8-203da76a9710 · outbound

This paper cites Cats and dogs.

One Last Attention for Your Vision-Language Model Cats and dogs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.350170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.350170Z digest=sha256:1d636685f618b00371e4636b9a7cd0a5ac3e716dab38d7d5335154c2bf9f6e4f

Observation 779e1ce5-b2f5-4364-ae70-0edef3638ac0 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

One Last Attention for Your Vision-Language Model Learn- ing transferable visual models from natural language super- vision

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.788954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.423283Z digest=sha256:7e9a72558f02f8c4f8ed203fe4ccd0d73e8e9857d05f69854e3a9bed939cd0d9

Observation 2f9248d8-b7b1-47f7-9218-c9d206f1f8e6 · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? InICML, 2019.

One Last Attention for Your Vision-Language Model Do imagenet classifiers generalize to im- agenet? InICML, 2019

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.428228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.428228Z digest=sha256:69204d38680b72a7b0cf0826d59b81f3d4f6b150b0befe7b67db05306426c1b5

Observation 2619913b-af4b-4c99-a640-00c443a5023c · outbound

This paper cites Towards parameter-efficient integration of pre- trained language models in temporal video grounding.

One Last Attention for Your Vision-Language Model Towards parameter-efficient integration of pre- trained language models in temporal video grounding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.778807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.430708Z digest=sha256:ec4fc860226a17443ec6f09b7562991ba60444281233da51c6cf86cced1b9da6

Observation fa70a0ce-9d68-4777-bb3e-bc0d626848ec · outbound

This paper cites Test- time prompt tuning for zero-shot generalization in vision- language models.

One Last Attention for Your Vision-Language Model Test- time prompt tuning for zero-shot generalization in vision- language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.772165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.432660Z digest=sha256:84dbc0202da05b2e4526561993a3339105ce18ac42f8fa7892024807d5c6585e

Observation cad5eda9-24c2-4941-a8ea-f89d0c988cf1 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

One Last Attention for Your Vision-Language Model UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.434722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.434722Z digest=sha256:ad8cd3da7e3cbb54f6854904737c5fdf70f7f1af42ebb998466ec4f8011cf7ec

Observation d23df06a-72b1-4b50-a336-8ed954083cb1 · outbound

This paper cites Test-time training with self- supervision for generalization under distribution shifts.

One Last Attention for Your Vision-Language Model Test-time training with self- supervision for generalization under distribution shifts

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.764962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.437649Z digest=sha256:34cb4413242509c624b8fe37074d6fcf8c4639cec625d8db1a98e89686a3cf01

Observation a5c4ec97-e1bc-4078-8841-06758dd210df · outbound

This paper cites Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks.

One Last Attention for Your Vision-Language Model Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.758464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.439970Z digest=sha256:ec523e512399c218162b163cefeac56774f5298bf8a31b22d97cbfd3b8414057

Observation 2f8b4646-4474-4279-8f55-9ed25583b46c · outbound

This paper cites The information bottleneck method.

One Last Attention for Your Vision-Language Model The information bottleneck method

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.441899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.441899Z digest=sha256:fa2606c62e35da54b2cc2c6747daa4fd4031cf279b19665c69963cec0f8f840b

Observation 79f71ec9-7d94-430d-81a9-ec98b2ea5d3d · outbound

This paper cites Visualizing data using t-sne.JMLR, 9(11), 2008.

One Last Attention for Your Vision-Language Model Visualizing data using t-sne.JMLR, 9(11), 2008

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.751715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.444249Z digest=sha256:931f296472b3cd1edbb614acdda39b325dba18b144740498adae2acdc0565fb5

Observation 31758730-5040-4e60-9897-046bdc88e252 · outbound

This paper cites Attention is all you need.

One Last Attention for Your Vision-Language Model Attention is all you need

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.745284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.446291Z digest=sha256:8d67052f5cec3ce0547701ff7fb5acf60b287a221edbfad187093877257021fa

Observation b23944ac-2d87-4188-9460-edcc38dc012a · outbound

This paper cites Tent: Fully test-time adaptation by entropy minimization.

One Last Attention for Your Vision-Language Model Tent: Fully test-time adaptation by entropy minimization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.738958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.448328Z digest=sha256:f42705a6d572876158129d56c34ee6d6124ada976f2595c9299e1be3d1a44f49

Observation 4e166775-ef4f-4b91-87ff-cbddc4eca21e · outbound

This paper cites Learning robust global representations by penalizing local predictive power.

One Last Attention for Your Vision-Language Model Learning robust global representations by penalizing local predictive power

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.450753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.450753Z digest=sha256:0ece7e6092c92b6aa3a7fd2f2560c4308b9be43cbd62025636359afaef0218f8

Observation 95dd0666-7e44-4b3e-bf21-8c59c4a10674 · outbound

This paper cites Robust fine-tuning of zero-shot models.

One Last Attention for Your Vision-Language Model Robust fine-tuning of zero-shot models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.728739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.452861Z digest=sha256:c68c3715915fb3b93e09103b3a0127fc4f0c7bd9a82a121d010cdadceec69976

Observation 22ddd1b0-7c36-4a93-bdb6-e26ff3c73e1e · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo.

One Last Attention for Your Vision-Language Model Sun database: Large-scale scene recognition from abbey to zoo

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.455470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.455470Z digest=sha256:a1c199c452b9f8f0f25aea7c685c9fe3f50d2b9e0a9ae5c8aeeb6f5b43d3948b

Observation b49a82fa-951f-48ef-9666-241c2e2cf0a3 · outbound

This paper cites Explicit inductive bias for transfer learning with convolutional net- works.

One Last Attention for Your Vision-Language Model Explicit inductive bias for transfer learning with convolutional net- works

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.716235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.457851Z digest=sha256:61aa502ce887c92d7435c570babac5205a9b27e2b6f184d940ae199db165c68c

Observation ba807956-e646-412e-875f-fdfec71eced4 · outbound

This paper cites Mma: Multi-modal adapter for vision-language models.

One Last Attention for Your Vision-Language Model Mma: Multi-modal adapter for vision-language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.710424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.460336Z digest=sha256:7a46ca1270d049e2493ee8c8cef66d9ab8958e1a3e17f4d3a942e6cb29fec296

Observation 83eff9e0-d053-41ee-86eb-3dab779de8c0 · outbound

This paper cites Visual- language prompt tuning with knowledge-guided context op- timization.

One Last Attention for Your Vision-Language Model Visual- language prompt tuning with knowledge-guided context op- timization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.704017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.462589Z digest=sha256:0fb2c1fe428cbcdab5558fe4b71c0b314db75e06cff55521345a384b4c38e823

Observation 05fd11fb-97eb-45ba-9795-134c26d756d9 · outbound

This paper cites Filip: Fine-grained interactive language-image pre-training.

One Last Attention for Your Vision-Language Model Filip: Fine-grained interactive language-image pre-training

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.697968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.464690Z digest=sha256:02a151a6f9f6803f3cf305f453485bd2251c6178a273a2857beeb5bb2485544c

Observation e04ae707-b228-4540-b38a-0bd58c63ec1b · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

One Last Attention for Your Vision-Language Model Florence: A New Foundation Model for Computer Vision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.466941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.466941Z digest=sha256:bb4f0d45f100402c83587d29e48453a81cc13e27bd0d6b8907801a38129a6e81

Observation 9ce889c6-3b7b-4414-9249-2591b49e9343 · outbound

This paper cites On the test-time zero- shot generalization of vision-language models: Do we really need prompt learning? InCVPR, 2024.

One Last Attention for Your Vision-Language Model On the test-time zero- shot generalization of vision-language models: Do we really need prompt learning? InCVPR, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.691779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.469568Z digest=sha256:a5c8ddb94b00d635a3e84b593f95e978ad15ef472b6f3e210074574b17ad2f80

Observation 4e1bf780-8215-4fab-ba9e-e3142fd70d09 · outbound

This paper cites Unified Vision and Language Prompt Learning.

One Last Attention for Your Vision-Language Model Unified Vision and Language Prompt Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.472212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.472212Z digest=sha256:a61553052bdd3ac497676cf463de143b8988674f2d40cf3e0353a6b7451e55b5

Observation 3f0159fe-460e-41d4-951d-032cbe362786 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

One Last Attention for Your Vision-Language Model Lit: Zero-shot transfer with locked-image text tuning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.685112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.474653Z digest=sha256:7c21fa6f92b72e5d4f0198065ca0a08e9f437dac03397a203e0689eb57196b69

Observation f11ace67-f3e9-4537-9b16-8c6aa389621e · outbound

This paper cites Sigmoid loss for language image pre-training.

One Last Attention for Your Vision-Language Model Sigmoid loss for language image pre-training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.678790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.477069Z digest=sha256:093c9ec10d3485bbd7fe2cfe88d4d2303229357d3cd2d025ff55435c315fbe76

Observation 10b35865-6f5b-4912-9a6a-262893b6697f · outbound

This paper cites Dept: Decoupled prompt tuning.

One Last Attention for Your Vision-Language Model Dept: Decoupled prompt tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.672157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.479244Z digest=sha256:30e707e8c96b7f97dcca3d1a0b77984640dec227223c54442dc46f33b4271428

Observation b16be42a-3d49-4ac0-bd57-2a0ff8eb4a0d · outbound

This paper cites Tip-adapter: Training-free clip-adapter for better vision- language modeling.

One Last Attention for Your Vision-Language Model Tip-adapter: Training-free clip-adapter for better vision- language modeling

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.665583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.481674Z digest=sha256:93f0e9915f07632ab964422cc899492308729f8edc48bfe55743d41061e69a66

Observation 4ad50816-c248-4367-a7d6-41b1f42c3860 · outbound

This paper cites Contrastive learning of medical visual representations from paired images and text.

One Last Attention for Your Vision-Language Model Contrastive learning of medical visual representations from paired images and text

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.658055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.484156Z digest=sha256:9673cb3101484dacf360c86c9d8593cd84d5e8109d4b277aaf9aaa61ed1f38e2

Observation 0b89c9b0-e712-437e-8243-2b5620fdf9ee · outbound

This paper cites Test-time adaptation with CLIP reward for zero-shot gener- alization in vision-language models.

One Last Attention for Your Vision-Language Model Test-time adaptation with CLIP reward for zero-shot gener- alization in vision-language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.651821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.486387Z digest=sha256:80e36f9470dbc0af026e72989ea7ecea4ccfb2057858b7bf4218817070772bb1

Observation 094ee6ae-acb3-4ee1-844f-fa96c536c8ce · outbound

This paper cites Conditional prompt learning for vision-language models.

One Last Attention for Your Vision-Language Model Conditional prompt learning for vision-language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.644663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.488483Z digest=sha256:3d28f0c272d2c1d59db33d41406aae72939837a71e587450f98d1ed61a810d10

Observation 103dbb33-df2f-4fcb-b776-e99d61f8419d · outbound

This paper cites Learning to prompt for vision-language models.IJCV, 130(9):2337–2348, 2022.

One Last Attention for Your Vision-Language Model Learning to prompt for vision-language models.IJCV, 130(9):2337–2348, 2022

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:42:16.637664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.490812Z digest=sha256:f4141433652244c9b454c0a1197b5ce172f7ef60680efb45faa3f9a37a77d348

Observation 573d9c7d-1b03-47f8-99ba-c30e38096a83 · outbound

This paper cites Prompt-aligned gradient for prompt tuning.

One Last Attention for Your Vision-Language Model Prompt-aligned gradient for prompt tuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:16.493608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:16.493608Z digest=sha256:f2c3225aa475d880f95badbf7fe7acb204c0bb79cb0a8c8f4a3f6cd8783058fd

Observation bbe56590-6381-4766-a7ea-2159f28bd6bb · outbound

This paper cites an unresolved cited work.

One Last Attention for Your Vision-Language Model Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:42:16.619436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.498525Z digest=sha256:865d6845ac3d7dd8325af990f57a2d8dda5dc50795e681ea9c6b53a508545f5a

Observation 3359f93b-f114-44d9-a5c1-434bc51df1ab · outbound

This paper cites an unresolved cited work.

One Last Attention for Your Vision-Language Model Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:42:16.612667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.501196Z digest=sha256:c7595444568595c63871f06522a4e7af4453109af16898b693d8fe52c94b54ac

Observation a7c747b0-b24a-4f85-ac34-562a4dadde03 · outbound

This paper cites an unresolved cited work.

One Last Attention for Your Vision-Language Model Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:42:16.605544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.503631Z digest=sha256:e06ea3704c142af57d47f6c6c6c0623e3e0b39176278b2e34540617a8679ccbf

Observation 23b32ef5-7b8d-43de-a7c6-b7ba33d02163 · outbound

This paper cites an unresolved cited work.

One Last Attention for Your Vision-Language Model Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:42:16.599082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.506226Z digest=sha256:fb7af658074198ab4dd266d2010716d051b1e1a82af6bb8fab261f2870995397

Observation ffd7b327-ec0f-407e-8e76-a9be7d8984ad · outbound

This paper cites an unresolved cited work.

One Last Attention for Your Vision-Language Model Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:42:16.592721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.508717Z digest=sha256:2cc2a335eee229e721555e543437a2a7de417bdb7e8fd1dcbd50a8ba25feabf9

Observation 9b1e13f8-f4e7-45d9-9509-873b823b639d · outbound

This paper cites an unresolved cited work.

One Last Attention for Your Vision-Language Model Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:42:16.586440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.510946Z digest=sha256:420e40ba828f2961bbafcd85c4110c92495fecba6287afb212c9e02bf3a51505

Observation af778d57-1ff4-4d95-b488-6d4e478ff2d8 · outbound

This paper cites an unresolved cited work.

One Last Attention for Your Vision-Language Model Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:42:16.580101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.513762Z digest=sha256:a36ba4ffef4f41d6e5128ee5c9e156c55b50b77ae8456edf6985b8dbfbf4eb06

Observation e4ec6e15-a9f9-4fca-afb6-33c5f6a5e5d1 · outbound

This paper cites an unresolved cited work.

One Last Attention for Your Vision-Language Model Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:42:16.573354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.515695Z digest=sha256:98fc18e601ef1ed40669f4b40b263a78cdc4dc8b8c8d7bbc0ab3cc2e42fc2680

Observation c9bf58b0-bae0-4806-bfdc-3dc927ec7d76 · outbound

This paper cites an unresolved cited work.

One Last Attention for Your Vision-Language Model Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:42:16.625722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T15:42:16.496076Z digest=sha256:0b76f030898c84055c0107021fd6662ae1f43167d5057f7581857a0c3662ce07

Pith citing papers

No inbound Pith citation observations are available.