Pith. sign in

Paper Citation Record · LEDGER

GLAD: Generalizable Tuning for Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 3 inbound Pith citation observations for arXiv:2507.13089.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13089 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:37:24.063644Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:24:53.437554Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact3
  • verified fuzzy44
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8c631083-f522-4e07-a307-b4355579d76a · outbound

This paper cites Food-101–mining discriminative components with random forests.

GLAD: Generalizable Tuning for Vision-Language Models Food-101–mining discriminative components with random forests

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:17.647963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:17.647963Z digest=sha256:34a6704059f62a4042f75902236a8f9ad2dec95f6a12d39e8028118a800fbe16

Observation 18c46557-4c87-4e5d-9b2b-23ae3baf775b · outbound

This paper cites Domain prompt learning with quaternion networks.

GLAD: Generalizable Tuning for Vision-Language Models Domain prompt learning with quaternion networks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.642311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:17.697297Z digest=sha256:c6ea018f6740fec9e630230a91d3053abb995cbeb6d5b4e523f0bda2bc0a8819

Observation f5650e68-a17e-43f8-802b-123d6870fa79 · outbound

This paper cites Tokenmixup: Efficient attention-guided token-level data augmentation for transformers.

GLAD: Generalizable Tuning for Vision-Language Models Tokenmixup: Efficient attention-guided token-level data augmentation for transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.627929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:17.771183Z digest=sha256:1883f053eeb7c5366aab62e80643fd6e6975d792154ea5c5078625ce0fe62b08

Observation 7ef1785a-36f0-40fd-9ea4-f5bb33dd768d · outbound

This paper cites Describing textures in the wild.

GLAD: Generalizable Tuning for Vision-Language Models Describing textures in the wild

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:17.939549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:17.939549Z digest=sha256:765ebaf71386b27e38b65485c02cdd05542f6da59bd15c256c8a7ac03bc12ed1

Observation 0ad76c32-0c90-486f-b8e4-3ea618b7fccf · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

GLAD: Generalizable Tuning for Vision-Language Models Imagenet: A large-scale hierarchical image database

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.603846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:18.028779Z digest=sha256:b2a1d26ffb04805b081a9965d83eaa163da4675542d81f7c0ed52e4628a6649f

Observation 92e68a92-edaa-4485-8bfe-43b4805b65ee · outbound

This paper cites Learning to prompt for open-vocabulary ob- ject detection with vision-language model.

GLAD: Generalizable Tuning for Vision-Language Models Learning to prompt for open-vocabulary ob- ject detection with vision-language model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.589929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:18.104529Z digest=sha256:441bcfb8ca9f7d5e1f6d85af8461e126400053c8cd910483fdcfc4ca332d3eaf

Observation 219ff661-2bad-498b-8dc5-b478deef58ff · outbound

This paper cites Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories.

GLAD: Generalizable Tuning for Vision-Language Models Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.574755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:18.176667Z digest=sha256:e59f6db9d38f003a1c32a30d5fd72daa703d26c5520428c124260dacd1199c9f

Observation 1cdbdb46-701f-4237-be09-1ffb4e5a7ad3 · outbound

This paper cites Prompt- det: Towards open-vocabulary detection using uncurated im- ages.

GLAD: Generalizable Tuning for Vision-Language Models Prompt- det: Towards open-vocabulary detection using uncurated im- ages

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.555382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:18.272795Z digest=sha256:49ef57b514be80c69811d4f083646950920e9e518c70ee00e6cb0a91b834c0bd

Observation 3123dc9a-92b4-4b59-a7cf-d4ce526c1dec · outbound

This paper cites Sharpness-Aware Minimization for Efficiently Improving Generalization.

GLAD: Generalizable Tuning for Vision-Language Models Sharpness-Aware Minimization for Efficiently Improving Generalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:18.401393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:18.401393Z digest=sha256:29994870464088e8590bda1efbfe4fb7d4adaf57c8d92bddc36bfa0c44177174

Observation 062cd90d-afa2-40bf-b4ed-a5ecf1e1fef6 · outbound

This paper cites CLIP-Adapter: Better Vision-Language Models with Feature Adapters.

GLAD: Generalizable Tuning for Vision-Language Models CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:18.530702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:18.530702Z digest=sha256:d5cc8cff0dab72745238f62c66dcb48b6b724a6dc186d4d9748b89773918d0b5

Observation c8cdf682-f7dd-47c4-8bd7-fa1559d24ea6 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

GLAD: Generalizable Tuning for Vision-Language Models Explaining and Harnessing Adversarial Examples

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:18.616605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:18.616605Z digest=sha256:87bdac6c7e6910c92dbcbbf7ce6779c41af1d6c001a42502a2ea4ea48e8b4b91

Observation c557da71-b43d-4aa4-ad2c-953260e28b21 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

GLAD: Generalizable Tuning for Vision-Language Models Open-vocabulary object detection via vision and language knowledge distillation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.540229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:18.662303Z digest=sha256:0497d0a4426a5dfa1bfb12b6544a0fb42f497adb1b41035c0900612b9bde4bb7

Observation 6d2940be-bf72-4e9f-98a7-d6b6929d039f · outbound

This paper cites Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification.

GLAD: Generalizable Tuning for Vision-Language Models Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.525518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:18.791568Z digest=sha256:7952b5887fda0cce8d7445b31fc588ec5cf63a85cc84d8c869f31bbaa59dff6d

Observation f44a092f-0253-48bb-a48b-226d4d514e5d · outbound

This paper cites The many faces of robust- ness: A critical analysis of out-of-distribution generalization.

GLAD: Generalizable Tuning for Vision-Language Models The many faces of robust- ness: A critical analysis of out-of-distribution generalization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.509559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:18.857488Z digest=sha256:6c088246c5edddc98a074e1152d962a7b9eff986d72cff4805801f03daa3a408

Observation 89a001d4-2771-45ac-b562-633982e981f0 · outbound

This paper cites Natural adversarial examples.

GLAD: Generalizable Tuning for Vision-Language Models Natural adversarial examples

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.495020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:18.963976Z digest=sha256:a87673698d9be1687780b7fe3c33e94e8ae86410558f9d2d38416981784e5e7f

Observation 6b185be9-37b5-46cb-a022-c7019faad37d · outbound

This paper cites Lora: Low-rank adaptation of large language models.

GLAD: Generalizable Tuning for Vision-Language Models Lora: Low-rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.090863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.090863Z digest=sha256:0177a213b5bd43b13326fceb5bbebb3e5f314e0b0222aee4b0452a470594e335

Observation 00d50d34-2784-4f23-aa36-adebb894f1a2 · outbound

This paper cites Learning a Better Initialization for Soft Prompts via Meta-Learning.

GLAD: Generalizable Tuning for Vision-Language Models Learning a Better Initialization for Soft Prompts via Meta-Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:37:24.853475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:19.243859Z digest=sha256:74b09c76695cc34ee12c4e63a6b5ccec0c342094fe7a105f7cd2bd57dd1a0c02

Observation 7f0920ef-a73e-4379-8f36-d74481668e8a · outbound

This paper cites Patching open-vocabulary models by interpolating weights.

GLAD: Generalizable Tuning for Vision-Language Models Patching open-vocabulary models by interpolating weights

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.347269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.347269Z digest=sha256:c360b7d3fed06eda64574a43052e409dab044da27b166c7060b7da603cb86054

Observation 5ff13c00-186d-45e9-97e9-fff6c9f2e3fe · outbound

This paper cites Vi- sual prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models Vi- sual prompt tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.448372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.448372Z digest=sha256:442be352c6a5803cd2037e1cb761d76ca28526491f711fbe37be6c6813cd3aa5

Observation adf384f3-9ef9-412d-9a1c-73dffa930e01 · outbound

This paper cites Maple: Multi-modal prompt learning.

GLAD: Generalizable Tuning for Vision-Language Models Maple: Multi-modal prompt learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.456352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:19.508974Z digest=sha256:896eaf53dda8311c2b8503cd9f8440f8e253bb25e09a2c43d307ac43467d18b7

Observation d58f73ca-6225-4079-8f74-8bc9337af873 · outbound

This paper cites Self-regulating prompts: Foundational model adaptation without forgetting.

GLAD: Generalizable Tuning for Vision-Language Models Self-regulating prompts: Foundational model adaptation without forgetting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.440487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:19.619850Z digest=sha256:90eef2a0bc9be7ac6bbb4e3acb8c92f2a0e785d1f6420f9e2b08e638be454dd7

Observation ec6763f4-2486-43f2-8d8a-cdec6d4795fd · outbound

This paper cites Co-mixup: Saliency guided joint mixup with super- modular diversity.

GLAD: Generalizable Tuning for Vision-Language Models Co-mixup: Saliency guided joint mixup with super- modular diversity

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.425528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:19.807655Z digest=sha256:264e6b0029856c7e0951f417a75884c2ca74cdf3a59f568040fae20bf39b6e2b

Observation 5b63c5e1-dc1a-4444-a1e4-dd0c43ae81e3 · outbound

This paper cites 3d object representations for fine-grained categorization.

GLAD: Generalizable Tuning for Vision-Language Models 3d object representations for fine-grained categorization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.982942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.982942Z digest=sha256:b4bd1dc74f5761e25d50087e9af7fe2b2ad8800c43423a6ef7eead84995bfa91

Observation b57eb404-5fac-42f4-ada8-4ea5bff4d2d0 · outbound

This paper cites Read-only prompt op- timization for vision-language few-shot learning.

GLAD: Generalizable Tuning for Vision-Language Models Read-only prompt op- timization for vision-language few-shot learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.396667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:20.056213Z digest=sha256:c73b4b28d575d3b7fb17c1eae08df1a9092976ab5898863d40098dbac7f8df48

Observation 8db438d4-ce38-4fc4-a3f7-f481a3744da0 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models The power of scale for parameter-efficient prompt tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.380993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:20.165333Z digest=sha256:2680abfa33f93b3dbf75f23ebff08276b6454822e6f00b6830abdf2b6db04862

Observation 6825f01f-c45e-4115-805c-0a6b5cea8238 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

GLAD: Generalizable Tuning for Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.278857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.278857Z digest=sha256:07a3525350fd47a5825a874d8c8562cd9c5d1e49545a3b0b7b2355c4bf7dcd2d

Observation 8d1fffa2-0021-4fa0-a42f-194812283d3a · outbound

This paper cites Promptkd: Unsupervised prompt distillation for vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Promptkd: Unsupervised prompt distillation for vision-language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.463882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.463882Z digest=sha256:7775484addb1bb5ad6a92b6d1c4ace18c787f61acbde2d94d0f10ff7b9630168

Observation 1db82ce7-f022-48a6-abd9-dca872383d73 · outbound

This paper cites Visual instruction tuning.

GLAD: Generalizable Tuning for Vision-Language Models Visual instruction tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.549960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.549960Z digest=sha256:ebb7e5d907da5681d64f05505ba0dc5579470e49a65426c106ab9ca36760363f

Observation a3d9f104-71f8-4166-86f3-088edfee8a7f · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing.

GLAD: Generalizable Tuning for Vision-Language Models Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.333469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:20.621651Z digest=sha256:f83497a6eef761acbe64335f04474be48068f822e1eada07806dd2fa71bb9155

Observation 9afda434-7a57-409d-b40c-f35ca7b1eb54 · outbound

This paper cites Decoupled weight decay regularization.

GLAD: Generalizable Tuning for Vision-Language Models Decoupled weight decay regularization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.318282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:20.716825Z digest=sha256:1463cfeeea9256a90853782f284ac6fe502a0708f8d2041f5e0d473a7290af96

Observation ba63cdc1-477a-44e0-b77a-aef4a4a45e5f · outbound

This paper cites Prompt distribution learning.

GLAD: Generalizable Tuning for Vision-Language Models Prompt distribution learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.302878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:20.801162Z digest=sha256:e1c3d308b693c459a7f2f819ed69be66540871a18c878d40fb583318a68c0920

Observation 24a22a1d-c2eb-4b73-ad51-a3ba60b337d0 · outbound

This paper cites Image segmentation using text and image prompts.

GLAD: Generalizable Tuning for Vision-Language Models Image segmentation using text and image prompts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.912376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.912376Z digest=sha256:2ca9b694830248be2fb5469818d53741a49085d4db15f412e287f1f004df7db8

Observation a25130ff-2c5c-44c8-8ab4-ba7bad27944f · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

GLAD: Generalizable Tuning for Vision-Language Models Fine-Grained Visual Classification of Aircraft

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:21.038250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:21.038250Z digest=sha256:d3afcbdb42ac69cc48dfc99e7beaa4b3a16c7980ba95ab3275ed2f75d893250f

Observation 5772af9d-23eb-449e-94e4-c3813760953a · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

GLAD: Generalizable Tuning for Vision-Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:21.189934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:21.189934Z digest=sha256:cdb7bfd2e6bbc4ff5a0654d1efc0274f823f9b4cc30d6b8a8c57ddfdd1b0e884

Observation 715a6811-a62d-4ac5-b41b-be757be0b828 · outbound

This paper cites Lookbehind-SAM: k steps back, 1 step forward.

GLAD: Generalizable Tuning for Vision-Language Models Lookbehind-SAM: k steps back, 1 step forward

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:37:24.559955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:21.263196Z digest=sha256:ca85a4cf7d72e3600399d94264b24f2fabc4a6671868e16912ca991392aeb097

Observation cea1fbff-c8c4-4478-bdc4-839e46f9c63a · outbound

This paper cites When does label smoothing help? Advances in neural in- formation processing systems, 32, 2019.

GLAD: Generalizable Tuning for Vision-Language Models When does label smoothing help? Advances in neural in- formation processing systems, 32, 2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.275633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:21.345862Z digest=sha256:0dd357b03700bcb923f101024382370cfa01b758a8dae4e64ac9cc469d7d4909

Observation cbe497ab-911b-43df-b53c-d286e638ee63 · outbound

This paper cites Automated flower classification over a large number of classes.

GLAD: Generalizable Tuning for Vision-Language Models Automated flower classification over a large number of classes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.255053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:21.409595Z digest=sha256:88d803a36417f0717f6a7a875027b8d1fc547421e0d5b03122c26655a81a6fd0

Observation 50e15440-8eba-4475-8bd7-813569dbe26d · outbound

This paper cites Metropolis-hastings data augmentation for graph neu- ral networks.

GLAD: Generalizable Tuning for Vision-Language Models Metropolis-hastings data augmentation for graph neu- ral networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.239720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:21.560898Z digest=sha256:9414648efc422bbee949ec278450b7af643b2e1daf9325a118e7215a9e8f4a1d

Observation 2dd8802d-63e3-42cc-8c26-ab2ba19371a3 · outbound

This paper cites Cats and dogs.

GLAD: Generalizable Tuning for Vision-Language Models Cats and dogs

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.224586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:21.659196Z digest=sha256:7db0fb67ae60ac85e104611d84cc8aeccb25f64cff82303349e22d4aa5f272b3

Observation 4efa1e0e-858b-4470-91e6-cb0d0e004511 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

GLAD: Generalizable Tuning for Vision-Language Models Learn- ing transferable visual models from natural language super- vision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.209884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:21.754468Z digest=sha256:39decaa3bb5ac1915c801b1a5b32f6f931a1dd6383b983cbe812cd2871ebcd57

Observation f66149b4-eb02-46e6-bea3-8b47a049627a · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400.

GLAD: Generalizable Tuning for Vision-Language Models Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.193331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:21.848212Z digest=sha256:b9008cecade8f55b79d2963c7139e02538b36a75ffdf22b7c38a580f58af7fc1

Observation 8ab4ee32-8f35-4cfc-b375-66f9e12e4cd6 · outbound

This paper cites Multimodal Instruction Tuning with Conditional Mixture of LoRA.

GLAD: Generalizable Tuning for Vision-Language Models Multimodal Instruction Tuning with Conditional Mixture of LoRA

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:37:24.375379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:21.969328Z digest=sha256:bed3291052ccecd8c4fe5630b554a1a86129d0f409d9c761bd56b9dfb69651fb

Observation 15bc974f-d69b-4c6c-bbac-a30160cdad02 · outbound

This paper cites Flava: A foundational language and vision alignment model.

GLAD: Generalizable Tuning for Vision-Language Models Flava: A foundational language and vision alignment model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.175590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:22.053563Z digest=sha256:2dbd0d6cdf85905eb7b40bf137b9f021003e82383f4b85d1ca07f902fd04f41c

Observation a6321269-a557-4ccf-98b8-5e6577c2a80b · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

GLAD: Generalizable Tuning for Vision-Language Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:22.126152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:22.126152Z digest=sha256:e838329df5e797f3103c0e1cace75a7ebe461abcdeda8565463acc92f8bb8a03

Observation 87838670-f2e1-4caa-8d19-a87cf28c806e · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.

GLAD: Generalizable Tuning for Vision-Language Models Dropout: a simple way to prevent neural networks from overfitting

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.158376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:22.207731Z digest=sha256:4aa2b0e514a7c5622566784c484851bdbc81011fd59dfe69d2d762411b3cda18

Observation 64b15ba4-b80f-4605-8475-87b9d93f4409 · outbound

This paper cites Rethinking the inception ar- chitecture for computer vision.

GLAD: Generalizable Tuning for Vision-Language Models Rethinking the inception ar- chitecture for computer vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:22.304601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:22.304601Z digest=sha256:8c96b15149403715f7540aa2a1445aa865c521c2c3c1e0b7c381374735521c92

Observation f35469ff-89a5-4ed7-97a5-5237913a7af8 · outbound

This paper cites Saliencymix: A saliency guided data augmentation strategy for better regularization.

GLAD: Generalizable Tuning for Vision-Language Models Saliencymix: A saliency guided data augmentation strategy for better regularization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.126994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:22.449012Z digest=sha256:517a7fe1514bc3555246504bb2b312c8b4da1614fe05e46927baed302d0e4ae1

Observation 19d8d1bc-bc71-4e05-ba88-830e49bcec3d · outbound

This paper cites Manifold mixup: learning better representations by in- terpolating hidden states.

GLAD: Generalizable Tuning for Vision-Language Models Manifold mixup: learning better representations by in- terpolating hidden states

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.113352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:22.547152Z digest=sha256:4c3e5c9cdec32dae0d460a38b89e8a8a5fc81d19779c531447eb837f9a357578

Observation c0af567c-bd48-4819-a7cd-1d224f61e30c · outbound

This paper cites Tuning multi-mode token- level prompt alignment across modalities.

GLAD: Generalizable Tuning for Vision-Language Models Tuning multi-mode token- level prompt alignment across modalities

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.098342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:22.642193Z digest=sha256:48f7af0cac1c6f0a44f75bdaa4d28d53fa7801d50a56000c530d6d0f91697739

Observation a9d23265-ea2c-46a4-b7c7-f71efadfa374 · outbound

This paper cites Learning robust global representations by penalizing local predictive power.

GLAD: Generalizable Tuning for Vision-Language Models Learning robust global representations by penalizing local predictive power

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.066014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:22.737119Z digest=sha256:92646ee6a39ab0664b5d97ccb6616db46dc6f4c6b9dc7e94ca7b597ebdf02725

Observation c7bb7d1a-3b7b-46e3-8776-df01c172b0cf · outbound

This paper cites Sharpness-aware gradient matching for domain generaliza- tion.

GLAD: Generalizable Tuning for Vision-Language Models Sharpness-aware gradient matching for domain generaliza- tion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.896307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:22.854498Z digest=sha256:576ab34afa0314d75d29f6b7fc111c43e112c1ead7c855c194d133e7a49b0e16

Observation 93c1dd62-3c84-4a36-bd81-82d3afda935f · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

GLAD: Generalizable Tuning for Vision-Language Models Cogvlm: Visual expert for pretrained language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.723422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:22.959808Z digest=sha256:f5fd077d3c9b623b6154898c239e2dad2a5edcd9c93d03804561818c4bfba9b8

Observation d3ea47da-5c28-4003-ba0e-dcabcfd057ed · outbound

This paper cites Robust fine-tuning of zero-shot models.

GLAD: Generalizable Tuning for Vision-Language Models Robust fine-tuning of zero-shot models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.583409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.033154Z digest=sha256:9a17b2784f35d4a552579201d0e407a1b0f38ffa57210928c825a7d97da48dd6

Observation 26048ba1-384a-4013-ab3e-dc289bd39204 · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo.

GLAD: Generalizable Tuning for Vision-Language Models Sun database: Large-scale scene recognition from abbey to zoo

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.403579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.079232Z digest=sha256:f589465ebddb8ccbfe653b11051c5815b55a11ce4697201bf499601a3a9d7b57

Observation 35131259-a89b-4284-bdad-790893de0a37 · outbound

This paper cites Tcp: Textual- based class-aware prompt tuning for visual-language model.

GLAD: Generalizable Tuning for Vision-Language Models Tcp: Textual- based class-aware prompt tuning for visual-language model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.246670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.150721Z digest=sha256:6e57f8182569cda0e1128c4d539df03739a2b44cc316eb927479f0f4bc29e921

Observation 3806585f-bbcc-4e70-be8d-2fde26295261 · outbound

This paper cites Cutmix: Regu- larization strategy to train strong classifiers with localizable features.

GLAD: Generalizable Tuning for Vision-Language Models Cutmix: Regu- larization strategy to train strong classifiers with localizable features

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.078177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.220423Z digest=sha256:037b689e4262afeed27c628dfd6cdf630bf6052a0c561e9b2a7d3bf8aeb8e192

Observation ebfd872d-fb0f-49e5-88fb-d918115b5910 · outbound

This paper cites Low-rank few-shot adaptation of vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Low-rank few-shot adaptation of vision-language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.915990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.292963Z digest=sha256:4942d5f8de5623b96b83b74837b238847714cbf8f769ba430a0d54ef6f514c64

Observation 626eb206-3728-4927-90aa-835e1c098562 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

GLAD: Generalizable Tuning for Vision-Language Models Lit: Zero-shot transfer with locked-image text tuning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.763594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.388218Z digest=sha256:8e49f0191f9b6131e9cfbe810678bff7ea7e202a7f236ef081460eb5033d29c7

Observation ed262ff5-852e-41f0-9a30-dbb356a79355 · outbound

This paper cites Three mechanisms of weight decay regularization.

GLAD: Generalizable Tuning for Vision-Language Models Three mechanisms of weight decay regularization

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.620181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.425869Z digest=sha256:9d25ce4232252dfc3f8b4d3f9ebdbaa3095566b963d959b4c0d9019cd9fb3d17

Observation d28e2228-19a9-496c-8ffc-62e3c36aeeb4 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

GLAD: Generalizable Tuning for Vision-Language Models mixup: Beyond Empirical Risk Minimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.459308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.459308Z digest=sha256:681ee33b6c5680388355e9f8512f1f434aee7e92c26d2aaea34882072a97c109

Observation 4834b420-d026-4e38-86e4-76b17aadee68 · outbound

This paper cites Dept: Decoupled prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models Dept: Decoupled prompt tuning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.426058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.526271Z digest=sha256:5a60e7f0cace679710166bab36fb536fd83ad8d2951ac2433b635e924261a70e

Observation a21fe1d1-d66b-4447-9f44-667a3af33382 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

GLAD: Generalizable Tuning for Vision-Language Models Adding conditional control to text-to-image diffusion models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.593178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.593178Z digest=sha256:028e9e55c3ff8e2169a426590740c734e31e99719a185bf19f6c124ef8e67b0b

Observation b7512f2a-6b89-4be1-b313-2e6847175b69 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of language models with zero-init attention.

GLAD: Generalizable Tuning for Vision-Language Models Llama-adapter: Efficient fine-tuning of language models with zero-init attention

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.126021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.638487Z digest=sha256:47db3a2c4a604d7938587dc60bce60a504b28e957f40ca49ba157bfe33a566cb

Observation c04e7633-c62d-43a3-ba52-725bb69e0a6d · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

GLAD: Generalizable Tuning for Vision-Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.687949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.687949Z digest=sha256:e195b10576c02a54f9ff17ed439310ed6eb08183aa083d447f57199d839ffbea

Observation 2dadc935-ceed-472a-a3cd-822b7c5036d8 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

GLAD: Generalizable Tuning for Vision-Language Models Regionclip: Region-based language-image pretraining

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:25.879634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.748806Z digest=sha256:1b8d12fb8dc01fe7958cfd031c466ebb04335d15b4757ad8daf87fe3f96c7b5f

Observation 7d993c5f-112d-4b83-b15e-97df888bc360 · outbound

This paper cites Conditional prompt learning for vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Conditional prompt learning for vision-language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:25.618806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.856095Z digest=sha256:c0ee45d943ba6109e08f8f53438a5df113a5efa70a95ff079f5fd2eae6a6d03c

Observation 2cbf78dd-d784-4289-a2ae-135dcd884f41 · outbound

This paper cites Learning to prompt for vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Learning to prompt for vision-language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:25.237348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:37:23.927447Z digest=sha256:dae6776297c989998fe940ba92239738113b80acaf1fa8fb713aea6c8c01834e

Observation e2a1921c-ccf8-4a8d-a30d-298f18332a4e · outbound

This paper cites Prompt-aligned gradient for prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models Prompt-aligned gradient for prompt tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.993559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.993559Z digest=sha256:4443efb96ed64558c17768668a9aac482019bf4e316ddb635e55a92ec28d6400

Observation e9c866b0-7b5d-4199-be01-49ece54bf2c4 · outbound

This paper cites Surrogate Gap Minimization Improves Sharpness-Aware Training.

GLAD: Generalizable Tuning for Vision-Language Models Surrogate Gap Minimization Improves Sharpness-Aware Training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:24.063644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:24.063644Z digest=sha256:516275942d97dd1032bcb1db9c4b7a2224ebf283c128e057e4e5f9d14fecc501

Pith citing papers

Observation bd78f140-02f5-4e16-b878-6d00bf653a39 · inbound

TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models cites this paper.

TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models GLAD: Generalizable Tuning for Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:53.437554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:53.437554Z digest=sha256:dc81cad69e357cfbc1cd489b34e629b588b3f7c49b1d8bc753a762c8da45acbf

Observation 30255c27-ef3b-4a30-b360-43510c56f66e · inbound

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models cites this paper.

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models GLAD: Generalizable Tuning for Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:40:26.018918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:39:15.820777Z digest=sha256:9d903c0ff511bedad410e174fde48be08307953831fdf83f1cc09e5568cfa7df

Observation 240109b8-82e3-4783-b64a-5230a8efe68a · inbound

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models cites this paper.

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models GLAD: Generalizable Tuning for Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T20:25:41.773345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:25:41.773345Z digest=sha256:61dae02835e970b551a5f707ba6e3f96c0c9d2ee2503e88fa77d092ea9f92022