Pith. sign in

Paper Citation Record · LEDGER

Context-Aware Multimodal Pretraining

As of 13 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 1 inbound Pith citation observation for arXiv:2411.15099.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15099 v1

Coverage vector

measured 100 of 113 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:37:57.214117Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:24:43.443062Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T19:24:44.625903Z

Reference resolution

100 of 113 outbound references displayed

  • verified exact1
  • verified fuzzy44
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70c765bd-bb09-4136-8981-ce3cf2be048d · outbound

This paper cites Towards in-context scene understanding.

Context-Aware Multimodal Pretraining Towards in-context scene understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.805098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.805098Z digest=sha256:7a29b609a5552fc63fddf7e281d3f6e3bcfd24682775e9522819bf2cbd9529e9

Observation c0e9a56c-d66c-4990-8c39-9eac67e6e79a · outbound

This paper cites Food-101 – mining discriminative components with ran- dom forests.

Context-Aware Multimodal Pretraining Food-101 – mining discriminative components with ran- dom forests

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.810558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.810558Z digest=sha256:a0518e104f809c2935c1b28ad652104499acb710c82af621fe56115690a5d197

Observation 77018c0c-9f9a-4993-8ebe-d7d729fea3ab · outbound

This paper cites JAX: composable transformations of Python+NumPy programs, 2018.

Context-Aware Multimodal Pretraining JAX: composable transformations of Python+NumPy programs, 2018

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.814909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.814909Z digest=sha256:d761fda299183276dd746e55f1d648be19237bc0297b3feb671593189777ae1d

Observation 45091425-3262-43a2-9787-d4ab8e48c3f3 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Context-Aware Multimodal Pretraining Emerging properties in self-supervised vision transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.818931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.818931Z digest=sha256:83149c61ce3dac70c1ce4fbb575bb1e78a4b5f7648c32bc5d1f044bf7b4bdc62

Observation 203c6aed-dd69-44da-872c-2cc3c2938ed6 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Context-Aware Multimodal Pretraining A simple framework for contrastive learning of visual representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.823242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.823242Z digest=sha256:a9ead17634e739edca3e9360ed7ff584adf3ca43ebbff3d27916720d0da2cc83

Observation 002d7e9c-d335-436e-8bfe-d192d5a83437 · outbound

This paper cites PaLI: A jointly- scaled multilingual language-image model.

Context-Aware Multimodal Pretraining PaLI: A jointly- scaled multilingual language-image model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.827291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.827291Z digest=sha256:46cefa41dd9ae4c87f6dff1d7e0f6869893c09d3dc285f8039274bc3e1f1d2d0

Observation 1983bc03-c6d4-4a5c-868e-40201bee3ca5 · outbound

This paper cites Meta-baseline: Exploring simple meta- learning for few-shot learning.

Context-Aware Multimodal Pretraining Meta-baseline: Exploring simple meta- learning for few-shot learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.831414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.831414Z digest=sha256:3a4f2b0964056f3babc36392b8a8d1f4e4f6ef930b415165c3116d4860df5607

Observation efd77a6e-def5-4984-8e05-27718a21572b · outbound

This paper cites Remote sens- ing image scene classification: Benchmark and state of the art.

Context-Aware Multimodal Pretraining Remote sens- ing image scene classification: Benchmark and state of the art

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.835669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.835669Z digest=sha256:b0d2925df711781191f7add2c46b2e3022330e014e2fcb2b8e47c409651a0f07

Observation 7cf14394-847a-4cfa-b7e4-d329de04a286 · outbound

This paper cites Cimpoi, S.

Context-Aware Multimodal Pretraining Cimpoi, S

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.839717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.839717Z digest=sha256:1ee7501fe3b3e76789760221a48aeaf5a7b3aef69bac616da9c464bcf312fe73

Observation caa77555-f1f3-4b56-a4f5-c5308f884a12 · outbound

This paper cites Embedding arithmetic of multi- modal queries for image retrieval.

Context-Aware Multimodal Pretraining Embedding arithmetic of multi- modal queries for image retrieval

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.843837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.843837Z digest=sha256:7ffdb94073af06d9316b69b9a07a0304d4f939de60712588ba611299ad89bc6e

Observation 0aeefcd9-4f7b-48a5-8438-b8a75a987473 · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.847839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.847839Z digest=sha256:86a040894ee671f2278882a484f65c8879384e32867e10496eb9c4dab4284e1c

Observation be71853a-f6b3-41e6-ac97-e522472987fa · outbound

This paper cites Calibrated Cache Model for Few-Shot Vision-Language Model Adaptation.

Context-Aware Multimodal Pretraining Calibrated Cache Model for Few-Shot Vision-Language Model Adaptation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.851770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.851770Z digest=sha256:520d23f06770f408e394e059ace22a47e3ac72b794f60c1e37adb71fd0ee6e1e

Observation 48e47d30-883a-4748-a93a-f4daee470184 · outbound

This paper cites An im- age is worth 16x16 words: Transformers for image recog- nition at scale.

Context-Aware Multimodal Pretraining An im- age is worth 16x16 words: Transformers for image recog- nition at scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.856394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.856394Z digest=sha256:49ae03adebe14b11fc7f2869923d30a0ef2a5654b109d0c72fcd187c89f4f280

Observation 791e29dc-2e3a-4abd-88d3-8197b9bbf257 · outbound

This paper cites With a little help from my friends: Nearest-neighbor contrastive learning of vi- sual representations.

Context-Aware Multimodal Pretraining With a little help from my friends: Nearest-neighbor contrastive learning of vi- sual representations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.860414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.860414Z digest=sha256:743bbb90b81e44dff850fac2336dbcb72bac57a07e3b96d557a6c4807796038e

Observation b1ab9afa-8d8e-4ce1-a16f-17109cceba86 · outbound

This paper cites Bad Students Make Great Teachers: Active Learning Accelerates Large-Scale Visual Understanding.

Context-Aware Multimodal Pretraining Bad Students Make Great Teachers: Active Learning Accelerates Large-Scale Visual Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.864270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.864270Z digest=sha256:e04a41d6f0338b2d5e371c8adac7edb22f70f925e5e454e7a08050734ac37b38

Observation 76b9f8e3-6223-46ff-b754-0bbe8ed0982f · outbound

This paper cites Data curation via joint example selection further accelerates multimodal learning.

Context-Aware Multimodal Pretraining Data curation via joint example selection further accelerates multimodal learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.868240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.868240Z digest=sha256:3f98eacfe558ed0c893b74b8892f94c1592482a2b118a56820b0563623529a57

Observation 1eebf32f-e751-4382-9dac-0216922ce997 · outbound

This paper cites Data determines distributional robustness in contrastive lan- guage image pre-training (clip).

Context-Aware Multimodal Pretraining Data determines distributional robustness in contrastive lan- guage image pre-training (clip)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.872532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.872532Z digest=sha256:44199628cc743bc0447997b197789f6934a2737b260d3a0bafefef145b516a83

Observation 3f82e4a1-ddc8-48d1-a892-5ea3aecd3922 · outbound

This paper cites Caption supervision enables robust learners.

Context-Aware Multimodal Pretraining Caption supervision enables robust learners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.876708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.876708Z digest=sha256:237e4d8cd3909b4d43acfb1fcfac711edb93f53bb94af925ce5b62d8d9bd9461

Observation 3344ca7d-cb0b-4469-958f-9d0f9eb12bf9 · outbound

This paper cites Context-aware meta-learning.

Context-Aware Multimodal Pretraining Context-aware meta-learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.881328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.881328Z digest=sha256:a1720b8f3d5e10f3efc6e7f6f83a07c705f9ff5b0d7c34b40ee9ae7d2d8f1450

Observation 3ecec771-1e71-44f9-9788-8068b4d7989e · outbound

This paper cites Model- agnostic meta-learning for fast adaptation of deep networks.

Context-Aware Multimodal Pretraining Model- agnostic meta-learning for fast adaptation of deep networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.885032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.885032Z digest=sha256:b6b3f5f7d3d9fb111f267d6a203da56d2bc1314c2d5accb52e0d72be052fbdef

Observation 4ce5867a-c163-46ac-90e3-90ae46b03a87 · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.

Context-Aware Multimodal Pretraining Clip-adapter: Better vision-language models with feature adapters

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.889441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.889441Z digest=sha256:9c5a92ea79b5d43893c068ca6e4eee30e1098f6b6fdc7605f4298dc08bdefc42

Observation ea69ba92-37d5-4ba2-ae99-f7e68baa6f06 · outbound

This paper cites Towards flexible perception with visual memory.

Context-Aware Multimodal Pretraining Towards flexible perception with visual memory

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.893039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.893039Z digest=sha256:93416c3259067b07edf07fed8269c0a05ad1155f075db0838be118843dbcaafe

Observation c1af1f92-3eac-4a0b-91d9-ba0b27715a8d · outbound

This paper cites Cyclip: Cyclic con- trastive language-image pretraining.

Context-Aware Multimodal Pretraining Cyclip: Cyclic con- trastive language-image pretraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.896970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.896970Z digest=sha256:fe889ea04a7c89ea7428c12f2708a574f9efc2b1f7ea298ef410a167f2db6997

Observation 9b34ef7e-5e95-4083-9bcb-6d08c7968367 · outbound

This paper cites kNN-CLIP: Retrieval enables training-free segmenta- tion on continually expanding large vocabularies.

Context-Aware Multimodal Pretraining kNN-CLIP: Retrieval enables training-free segmenta- tion on continually expanding large vocabularies

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.900622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.900622Z digest=sha256:516b9be8d385f5eb81e208817986f6ee9e69952c5615ee96e5f686e5cf74feb8

Observation 72718a6f-1246-4ebd-941a-9664a307b1c3 · outbound

This paper cites Calip: zero-shot enhancement of clip with parameter-free attention.

Context-Aware Multimodal Pretraining Calip: zero-shot enhancement of clip with parameter-free attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.904118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.904118Z digest=sha256:23b14764741f8e0ac9f3ce705ea1f5131f245b02e0079363aba5cb18fb4464aa

Observation c3d15297-de6e-461b-895b-ae9edcc6fb1e · outbound

This paper cites Anchor-based robust finetuning of vision-language models.

Context-Aware Multimodal Pretraining Anchor-based robust finetuning of vision-language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.907731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.907731Z digest=sha256:16fb192423542acc10870574e96afe41a92fc4bee9b0a65d7cd5bb231c9a760a

Observation 73352938-bd53-4430-858a-da6a8b60f17d · outbound

This paper cites Dota: Dis- tributional test-time adaptation of vision-language models.

Context-Aware Multimodal Pretraining Dota: Dis- tributional test-time adaptation of vision-language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.911882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.911882Z digest=sha256:cd578b44ba21cbbde2039a14eeb8228bccec61530d2c18e1c59f2a8ac791c1e8

Observation 31521a41-7721-487b-b3dc-687f24777925 · outbound

This paper cites Momentum Contrast for Unsupervised Visual Representation Learning.

Context-Aware Multimodal Pretraining Momentum Contrast for Unsupervised Visual Representation Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.915601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.915601Z digest=sha256:54cf5669a268c6161ff7b968be23ae1cfbd953c084911d4f92201983962220d8

Observation 04fa7485-abd6-4db2-843f-8b06ca16ce79 · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.

Context-Aware Multimodal Pretraining Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.919978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.919978Z digest=sha256:458081ee4d29dd1c9652dd36b560978340c99d0df9cc2d5509f54d9cb1a429c3

Observation 0f11c2db-bca9-4a8d-a902-59f97392e053 · outbound

This paper cites Ross, and Alireza Fathi.

Context-Aware Multimodal Pretraining Ross, and Alireza Fathi

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.923750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.923750Z digest=sha256:96a816a96677f4bd46abd1cbfda89fc535a58f7c82be0f4c7d0f1e66be2990d1

Observation 712c3c31-e308-433f-b43e-98f09a666d95 · outbound

This paper cites An open access repository of images on plant health to enable the development of mobile disease diagnostics.

Context-Aware Multimodal Pretraining An open access repository of images on plant health to enable the development of mobile disease diagnostics

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.927566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.927566Z digest=sha256:b37d2bc301e957b460f6c67ced74415b2b0efab66b4fe2709ba6504c911cbcab

Observation 65be5d2f-e5e3-4871-8b7c-c4c697f0b0cc · outbound

This paper cites Retrieval-enhanced contrastive vision-text mod- els.

Context-Aware Multimodal Pretraining Retrieval-enhanced contrastive vision-text mod- els

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.932485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.932485Z digest=sha256:443a50d75ecd1c2056b8b8dd5bfc06d03d92295ad7fb7413c7675ce4731ad24c

Observation 630f52a3-f4d3-4bba-bc2c-aabb5916cdb2 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Context-Aware Multimodal Pretraining Scaling up visual and vision-language representation learning with noisy text supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.936218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.936218Z digest=sha256:a39574380bd701833fe47c2dbe71ff77bafa00b92d6130acc2304dc6fbb68474

Observation 301c0320-f0e9-4ad3-927e-e8c3da1ea784 · outbound

This paper cites Billion- scale similarity search with gpus.

Context-Aware Multimodal Pretraining Billion- scale similarity search with gpus

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.939932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.939932Z digest=sha256:bc7c02318b51849fbd040978c69eaa16b59370a15e89b2be87d204e21dfd25f3

Observation 4e7093eb-6566-4511-bab5-8903fb954295 · outbound

This paper cites Multi-class texture analysis in colorectal cancer histology.

Context-Aware Multimodal Pretraining Multi-class texture analysis in colorectal cancer histology

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.943995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.943995Z digest=sha256:9cbebdc65a1f1d2a0acff17e3913764b1a7491741db35cc7568bbc36081ea2f1

Observation a12c991e-ddef-4b57-b744-3b4c493c9bdd · outbound

This paper cites Maple: Multi-modal prompt learning.

Context-Aware Multimodal Pretraining Maple: Multi-modal prompt learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.948744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.948744Z digest=sha256:eb71254bbcd342609c5ffdf36d36b3f3d06ed68fc5764ea1113cde8ec9f38b02

Observation 3f0178d9-5fb9-4878-9322-7773ebf4b246 · outbound

This paper cites Self-regulating prompts: Foundational model adaptation without forgetting.

Context-Aware Multimodal Pretraining Self-regulating prompts: Foundational model adaptation without forgetting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.953882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.953882Z digest=sha256:c07bfa8541dc31727bbcac60e4decee4f72f94d9483bc06e94c620d7da980741

Observation 40bbfb34-0c1b-4bb5-9f79-d237c8071b8b · outbound

This paper cites Datadream: Few-shot guided dataset generation.

Context-Aware Multimodal Pretraining Datadream: Few-shot guided dataset generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.958180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.958180Z digest=sha256:105b8f1475337a248f0bfb1efcf5e48fdc4ce36884ce62483271beaed278bd81

Observation 8083119b-cd98-48a2-8670-f63045aa784d · outbound

This paper cites Kirchhof, K.

Context-Aware Multimodal Pretraining Kirchhof, K

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.961933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.961933Z digest=sha256:32a975eb70624b3f9c553439967e6562cce4c7054a727bd50925ac70cd94a7c5

Observation 21f16633-77be-402c-b7ee-a9f2837d006b · outbound

This paper cites Wilds: A benchmark of in-the-wild distri- bution shifts.

Context-Aware Multimodal Pretraining Wilds: A benchmark of in-the-wild distri- bution shifts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.965448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.965448Z digest=sha256:7b33da87da09d7c3cb28fcd66ac88b56da92cf1f98f677b629272d8fc5f48ea5

Observation 52bea425-8108-4dfb-bd98-0a052fa4d840 · outbound

This paper cites 3d object representations for fine-grained categorization.

Context-Aware Multimodal Pretraining 3d object representations for fine-grained categorization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.969771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.969771Z digest=sha256:666d823a45cc31cd7e0b4ef82a5534b89f760973b1192fc169acc3da496ee7bb

Observation a56237f0-2a9a-4d2b-b657-7fa37b0660a4 · outbound

This paper cites Learning multiple layers of features from tiny images.

Context-Aware Multimodal Pretraining Learning multiple layers of features from tiny images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.346155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:56.973726Z digest=sha256:03a1a376fb3680a78a244fa18fe346ea4017cdee7b7f6d4a661d122b97135c78

Observation b070a9d6-f5af-4f61-91b8-64f339a3df6e · outbound

This paper cites SentencePiece: A sim- ple and language independent subword tokenizer and deto- kenizer for neural text processing.

Context-Aware Multimodal Pretraining SentencePiece: A sim- ple and language independent subword tokenizer and deto- kenizer for neural text processing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.335232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:56.977722Z digest=sha256:a0ca4ed9c81e4a85a1d635fe142d5d335cf22e3ce31dd384c36a93108c02f217

Observation 84fac99f-09e7-4b1c-99ad-abc348ab5315 · outbound

This paper cites Meta-learning with differentiable con- vex optimization.

Context-Aware Multimodal Pretraining Meta-learning with differentiable con- vex optimization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.323606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:56.981678Z digest=sha256:44afabda41caa6a34d42af961fed2c056eeadc64ab99f3335ce218c03ab27ba3

Observation 05d17675-4868-4313-be29-a586dc535122 · outbound

This paper cites Universal representation learning from multiple domains for few- shot classification.

Context-Aware Multimodal Pretraining Universal representation learning from multiple domains for few- shot classification

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.310950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:56.986973Z digest=sha256:0a30246b76357c26a9ac9148a7b35c23b459aba53f1a6b93d7c54c5e246ca8e2

Observation cc5f9c01-79bb-4ce3-b473-f5a0519ef951 · outbound

This paper cites The Devil is in the Few Shots: Iterative Visual Knowledge Completion for Few-shot Learning.

Context-Aware Multimodal Pretraining The Devil is in the Few Shots: Iterative Visual Knowledge Completion for Few-shot Learning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:37:57.462781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:56.991378Z digest=sha256:d563f47e57993c2bf04462fb6903b6571ec97b557b579a812280cb740e40fafe

Observation 100e03d1-66a3-4846-938d-ae726ac4617d · outbound

This paper cites Decoupled weight de- cay regularization.

Context-Aware Multimodal Pretraining Decoupled weight de- cay regularization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.299402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:56.995575Z digest=sha256:b243249fe412cb37403498615b1a8ad89d34566fa22caf0555d65472d78f84a7

Observation 64fc52eb-12c5-477a-8321-317b216e4e17 · outbound

This paper cites A closer look at few-shot classification again.

Context-Aware Multimodal Pretraining A closer look at few-shot classification again

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.287880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:56.999335Z digest=sha256:3a9a620a116a378492583856352bcacdd9042022c8441dedba3cffb67ac4ba96

Observation 66be19f4-ae75-41fa-a267-c0469a025420 · outbound

This paper cites Efficient and ro- bust approximate nearest neighbor search using hierarchi- cal navigable small world graphs.

Context-Aware Multimodal Pretraining Efficient and ro- bust approximate nearest neighbor search using hierarchi- cal navigable small world graphs

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.274989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.004285Z digest=sha256:67b2f83b91489a890fa56f59104ee1e4030b81c89bbd2383bec2211d084fa3ff

Observation 4c316190-2606-4644-8e32-d02c12dc136c · outbound

This paper cites Visual classification via description from large language models.

Context-Aware Multimodal Pretraining Visual classification via description from large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.262648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.008287Z digest=sha256:bbccb47a1b63484e61232f93e62fc8e2114e514506e1f4b533be431277d185b4

Observation 1c6ae813-cfd4-497c-915c-53081821f268 · outbound

This paper cites Understanding retrieval- augmented task adaptation for vision-language models.

Context-Aware Multimodal Pretraining Understanding retrieval- augmented task adaptation for vision-language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.249862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.012415Z digest=sha256:81a6ebd4c2122f1382f49c3d54399971f4f47d788d892cd24bf0aafbc078cf8a

Observation 64f891d0-f16e-4700-9fda-89deececb47a · outbound

This paper cites Slip: Self-supervision meets language-image pre- training.

Context-Aware Multimodal Pretraining Slip: Self-supervision meets language-image pre- training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.238857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.016068Z digest=sha256:371c32d9fa8df2b734090f43017e5f04b7f0d4dccb5d0bc816438dca72e6c910

Observation 34267b4b-0cd9-4537-88cb-0d25cf3faaee · outbound

This paper cites iCassava 2019 Fine-Grained Visual Categorization Challenge.

Context-Aware Multimodal Pretraining iCassava 2019 Fine-Grained Visual Categorization Challenge

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.019995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.019995Z digest=sha256:563a57b3e18814159911aec874dceb0e89f64e485d4448c827f647238ff73a86

Observation 6a6f778a-d042-4d47-b99e-0728aea139c9 · outbound

This paper cites Revisiting knn- based image classification system with high-capacity stor- age.

Context-Aware Multimodal Pretraining Revisiting knn- based image classification system with high-capacity stor- age

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.227426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.023854Z digest=sha256:cadd9b4009a81052975ce0fa80b2883478dbbfc6014ae93b096fb4c5a6088296

Observation a4ba33e8-925d-4752-8028-e3c9221f3ba8 · outbound

This paper cites On First-Order Meta-Learning Algorithms.

Context-Aware Multimodal Pretraining On First-Order Meta-Learning Algorithms

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.031237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.031237Z digest=sha256:3b58f270cc9da863eef598a3531f1a349ba8d63b31fe426bdb4ec22193d0f986

Observation a516c8b1-56ff-49db-ac39-ed9e9667bbbc · outbound

This paper cites CHiLS: Zero-shot image classifica- tion with hierarchical label sets.

Context-Aware Multimodal Pretraining CHiLS: Zero-shot image classifica- tion with hierarchical label sets

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.202661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.035558Z digest=sha256:094bf586865c621533d632670328b9d652bb0c358d50dd4e0415ef31df311d88

Observation ffc3d409-feb7-4a51-9ff2-9f8763f5a71b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Context-Aware Multimodal Pretraining Representation Learning with Contrastive Predictive Coding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.039478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.039478Z digest=sha256:f4afbb5200d66f280f5924c0c7c240ea20b7ba6fbddb7a4ba98776905414d457

Observation 442b0a8e-c71a-4a6a-91e7-640c56d7afa1 · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:58.191074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.043705Z digest=sha256:a4eb32e1c66622cae431818c41e4d5a182770419aff6c882388f80cb83e29020

Observation b27edba5-7226-4014-bcc7-df36401432a1 · outbound

This paper cites Svl-adapter: Self-supervised adapter for vision-language pretrained models.

Context-Aware Multimodal Pretraining Svl-adapter: Self-supervised adapter for vision-language pretrained models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.179609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.047926Z digest=sha256:3bc4c48f34e0f6ef5db0341137f1b9477e98574d53cbba5e3f4083fb38418024

Observation aad77fb2-7711-4292-a136-f79e59bc68a7 · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:58.167209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.051654Z digest=sha256:d8616c206ced5964c714f4d285537662c2fe2a282dc578c57a2de9fc6bfa74a8

Observation fb1fd241-37be-4acc-b86f-2b62a2cce57a · outbound

This paper cites Moment matching for multi-source domain adaptation.

Context-Aware Multimodal Pretraining Moment matching for multi-source domain adaptation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.154247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.056471Z digest=sha256:1662f3f24711a709c276c118fe026a7e6bbd3663ff5ac9c1e7324ecb27688e32

Observation 711b7eb8-5a14-40ff-825d-1fee0f1c1dbc · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:58.142966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.060344Z digest=sha256:22bbaeb2bc532191844b535acf9aa089b5b2857eb922ed1852890b4e7d7ece83

Observation a5a555e5-f8dc-4cf2-9f8a-6f8744cc603f · outbound

This paper cites Online Continual Learning Without the Storage Constraint.

Context-Aware Multimodal Pretraining Online Continual Learning Without the Storage Constraint

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.064283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.064283Z digest=sha256:161cef0b77de8472a6470d94b3f1f4dd9dfae28f4ad593fe0f41835e49e8e07c

Observation e006fed8-7986-4efa-a2df-edc5dbae5f72 · outbound

This paper cites What does a platypus look like? generating customized prompts for zero-shot image classification.

Context-Aware Multimodal Pretraining What does a platypus look like? generating customized prompts for zero-shot image classification

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.131420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.068671Z digest=sha256:cd74d9ed91cc06d36ab1f6c3e9a29371f409bcf391b6f72f9c0e31bee45120e7

Observation 9a24563c-aeb9-4cdb-9f12-218f0817d7ea · outbound

This paper cites Learning transferable visual models from natural language supervision.

Context-Aware Multimodal Pretraining Learning transferable visual models from natural language supervision

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.119743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.072389Z digest=sha256:b786870a74f0de7f8b92018b7e21147030b5dce72965f0c21ec75d57ce31db53

Observation 17366306-0fec-4a1c-bb20-4dd743293929 · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:58.108238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.076389Z digest=sha256:1ca1360e1a03b8149e4834ca3b0c0911cb82d26fe789d4cb21864fbaa6ecfaee

Observation 11a6dafd-f568-480b-8ea3-4b1e87e4bfd6 · outbound

This paper cites Meta-learning with implicit gradients.

Context-Aware Multimodal Pretraining Meta-learning with implicit gradients

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.097235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.080772Z digest=sha256:7adf62e46ad0424c7e268093568329642726e9b72b6da0e9c524fe675cd983a1

Observation 5a3eec50-6345-4d87-9c7a-e59d352db629 · outbound

This paper cites Towards to- tal recall in industrial anomaly detection.

Context-Aware Multimodal Pretraining Towards to- tal recall in industrial anomaly detection

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.086217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.084489Z digest=sha256:ffceb0981537000f99fa0e39127771447a09e5a61814a45cd14dd8438a0588fe

Observation f5547801-b058-45c5-b9aa-41767054ec72 · outbound

This paper cites Sophia Koepke, Oriol Vinyals, Cordelia Schmid, and Zeynep Akata.

Context-Aware Multimodal Pretraining Sophia Koepke, Oriol Vinyals, Cordelia Schmid, and Zeynep Akata

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.075211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.088022Z digest=sha256:422436f17bb0c8d0620bcdd045763505e67d1f8559059cda24816982845b581c

Observation e2ce42b8-6507-4191-bd56-8497b31970a0 · outbound

This paper cites A Practitioner's Guide to Continual Multimodal Pretraining.

Context-Aware Multimodal Pretraining A Practitioner's Guide to Continual Multimodal Pretraining

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.092746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.092746Z digest=sha256:71dd5962e413c9b24eaa100b236bde895ca28654641c9739c97e27594c78544a

Observation 03000239-5940-4e84-a6ab-09e07e705872 · outbound

This paper cites Berg, and Li Fei-Fei.

Context-Aware Multimodal Pretraining Berg, and Li Fei-Fei

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.063806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.097341Z digest=sha256:6a364b9206fe7b4dbef367a340b875ded5e1ef3ea2b6ce2aca08f16feee01e5f

Observation d6ecb8b8-8445-4d45-9567-e9ac8637df2e · outbound

This paper cites Rusu, Dushyant Rao, Jakub Sygnowski, Oriol Vinyals, Razvan Pascanu, Simon Osindero, and Raia Had- sell.

Context-Aware Multimodal Pretraining Rusu, Dushyant Rao, Jakub Sygnowski, Oriol Vinyals, Razvan Pascanu, Simon Osindero, and Raia Had- sell

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.052211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.100964Z digest=sha256:8eda0e7999d571848c99cd367f29a532863d063eb5c61e40c927a70def892d6c

Observation f2d45028-895b-4b9e-93e3-9e2990279037 · outbound

This paper cites Is a caption worth a thou- sand images? a study on representation learning.

Context-Aware Multimodal Pretraining Is a caption worth a thou- sand images? a study on representation learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.041055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.104716Z digest=sha256:e98f6ff91864b0ea14a9d221ee03883d35a7b5de8fb16c4ef68f54e92dbe3bef

Observation d93cf45e-2090-402e-95ef-4a92bd597ee6 · outbound

This paper cites Scott, Andrew C.

Context-Aware Multimodal Pretraining Scott, Andrew C

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.030097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.109119Z digest=sha256:cec13ea28cc65b4414ea1421241ddf3555775e9b23627f9e40ae286d2f4aaf0c

Observation c9b90e25-f56c-4b06-9818-850fc62b1c62 · outbound

This paper cites statsmodels: Econo- metric and statistical modeling with python.

Context-Aware Multimodal Pretraining statsmodels: Econo- metric and statistical modeling with python

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.012305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.113383Z digest=sha256:6dc58171d56b5036788e9e82a6e647b43fbc1e4634df9fbbbad3ff63c2b73486

Observation 81001eac-2be7-4ef4-a53c-7430a15c25a1 · outbound

This paper cites Prototyp- ical networks for few-shot learning.

Context-Aware Multimodal Pretraining Prototyp- ical networks for few-shot learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.999678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.117264Z digest=sha256:09486321d87177085509edc0e8669b80a8090e40ec95641c6766158e82bcdafe

Observation b477796c-5448-4111-9aaf-d90044714688 · outbound

This paper cites CLIP models are few-shot learners: Empirical stud- ies on VQA and visual entailment.

Context-Aware Multimodal Pretraining CLIP models are few-shot learners: Empirical stud- ies on VQA and visual entailment

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.976059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.124616Z digest=sha256:9fd502bc58fa881aa65553be65ab9bb5485a72218ee7d1a0ebb597ef429b005c

Observation 736c5b70-6251-4dd3-aa01-ad1ea67a5694 · outbound

This paper cites Momentum-based Weight Interpolation of Strong Zero-Shot Models for Continual Learning.

Context-Aware Multimodal Pretraining Momentum-based Weight Interpolation of Strong Zero-Shot Models for Continual Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.128358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.128358Z digest=sha256:c5d81acaaa57068e27b44a9ace794fbaf5b27adfcd91c868f4bb45f08726a939

Observation 007a8fe0-30bc-4ac9-9eae-401bb59f09b5 · outbound

This paper cites Torr, and Timothy M.

Context-Aware Multimodal Pretraining Torr, and Timothy M

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.964324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.132369Z digest=sha256:3ce2a61eec1e9e5e4c9ac1f1107efa39d6246e40d37856cac7697eafbbb36d5c

Observation 95c8b417-a95c-409d-8def-04ac52074166 · outbound

This paper cites A Fistful of Words: Learning Transferable Visual Models from Bag-of-Words Supervision.

Context-Aware Multimodal Pretraining A Fistful of Words: Learning Transferable Visual Models from Bag-of-Words Supervision

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.136457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.136457Z digest=sha256:bb3953aa2c7b8d282b61903c353753df914fbb4da6752e4bca208721d09bd3ee

Observation dc99c10a-e7c6-494d-817a-f680068153a7 · outbound

This paper cites Reflecting on the state of rehearsal-free continual learning with pretrained models.

Context-Aware Multimodal Pretraining Reflecting on the state of rehearsal-free continual learning with pretrained models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.140969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.140969Z digest=sha256:2c081bc8c3b139324462240743956c20b255561a2ff34b7d8637b55489c75175

Observation 2f283015-664d-4491-be18-b0dd2ff33416 · outbound

This paper cites Tenenbaum, and Phillip Isola.

Context-Aware Multimodal Pretraining Tenenbaum, and Phillip Isola

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.952387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.145029Z digest=sha256:205ba1444d83001bbcaca7f05c128f44742786480fbbd461645f0d6594810b25

Observation 723165a3-c8b8-48bf-90e5-3e7bd85928a8 · outbound

This paper cites Learning a universal template for few-shot dataset generalization.

Context-Aware Multimodal Pretraining Learning a universal template for few-shot dataset generalization

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.939953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.149162Z digest=sha256:9cd26a1b0e0bcb7ef25e39dda41fe68b7acacaf276ad453a8bd431998d121bbd

Observation 3ab390aa-eb50-4dae-a438-3083e747ee7e · outbound

This paper cites Sus-x: Training-free name-only transfer of vision-language models.

Context-Aware Multimodal Pretraining Sus-x: Training-free name-only transfer of vision-language models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.926951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.152895Z digest=sha256:fead12a67e895546776a1f2fefc8f06c3b184cbc7454fee7b50c2409ed4c09d6

Observation dea92e67-8c29-41a8-877e-579c86f2bb22 · outbound

This paper cites No ”zero-shot” without exponential data: Pretraining concept frequency determines multimodal model performance.

Context-Aware Multimodal Pretraining No ”zero-shot” without exponential data: Pretraining concept frequency determines multimodal model performance

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.914761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.156538Z digest=sha256:a38736252ea2728471a78b45c2f9a53676c5a6e3ada4f0ea49e566129df3b014

Observation 82c60407-9136-4ee8-bcb4-085c2e53757a · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Context-Aware Multimodal Pretraining Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.903118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.160216Z digest=sha256:8408a3e85e04a795a9d611985007aa01a236b5d7b75cbca7779ed0f2e1e3623b

Observation 9d0d702b-5736-4283-8759-998f518144f6 · outbound

This paper cites Matching networks for one shot learning.

Context-Aware Multimodal Pretraining Matching networks for one shot learning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.891598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.164036Z digest=sha256:1e8ba8f17cbe10d9ce7dbb465535b51aaa880265d374b4b2613fb394f252f63b

Observation 79ea8769-f9a9-4ad6-932c-235425f2753b · outbound

This paper cites Learning robust global representations by penalizing local predictive power.

Context-Aware Multimodal Pretraining Learning robust global representations by penalizing local predictive power

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.879394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.167806Z digest=sha256:d597988ce4a2c46696871aa921a50559a85f5cfa533df2adb1f43a8362742c97

Observation 58d52a3f-8c90-4eb9-ba33-21098263e4b3 · outbound

This paper cites SimpleShot: Revisiting Nearest-Neighbor Classification for Few-Shot Learning.

Context-Aware Multimodal Pretraining SimpleShot: Revisiting Nearest-Neighbor Classification for Few-Shot Learning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.172045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.172045Z digest=sha256:23950884b473ddd9f3f282fb147461d9a568423adda2372908f052ecd8854440

Observation 51f8e675-7350-412a-94a5-0b679013241e · outbound

This paper cites A hard-to-beat baseline for training- free CLIP-based adaptation.

Context-Aware Multimodal Pretraining A hard-to-beat baseline for training- free CLIP-based adaptation

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.868102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.175826Z digest=sha256:cc1cac5aa8c0f4aac9bc87999aa4a7965194a753bbf8512024f3b8a895b2d368

Observation 181e1fb5-0c3b-47f4-b543-a17c2dd03b85 · outbound

This paper cites Welinder, S.

Context-Aware Multimodal Pretraining Welinder, S

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.857105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.179462Z digest=sha256:85e20e881c0ddd564dadf65addd013fd7d7db905cb31df3b7c2fb4d368430cf3

Observation c99e5433-c297-4b5d-ba9b-e6f829b23660 · outbound

This paper cites Cascade prompt learning for vision-language model adaptation.

Context-Aware Multimodal Pretraining Cascade prompt learning for vision-language model adaptation

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.846173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.183149Z digest=sha256:ea4ccd535e6e775493314b3ebbb8acf9a99e02e1d3f58d81f5a942a302e2075c

Observation 2ef29238-f14e-4282-954b-6baa244b04ed · outbound

This paper cites Yu, and Dahua Lin.

Context-Aware Multimodal Pretraining Yu, and Dahua Lin

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.835015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.186746Z digest=sha256:a3b5f970e6d1af73650fab20f69a6a25205e69afb99853f4f5c375449180dd28

Observation deb49dfa-4408-4827-811a-446eb5d7debb · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:57.824058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.190332Z digest=sha256:69fc1b46a5ea37ee52b9b4da3c4ceb708f52bf78261f93297f33abeccc58afe3

Observation 575cea0f-38ef-425c-a187-745270fb36a0 · outbound

This paper cites Ra-clip: Retrieval augmented contrastive language-image pre-training.

Context-Aware Multimodal Pretraining Ra-clip: Retrieval augmented contrastive language-image pre-training

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.812976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.193803Z digest=sha256:08cc9959fe689a1fd0a004fa2a415dacdc219c34c41e6b41d339ab51c035ec9b

Observation e67708da-7c0d-4613-a188-06e4cdf8e229 · outbound

This paper cites MetaFun: Meta-learning with itera- tive functional updates.

Context-Aware Multimodal Pretraining MetaFun: Meta-learning with itera- tive functional updates

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.801712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.197717Z digest=sha256:ff3393f81f96e63c44f833ae9d6005d9aa2db6c3e2b30ad7a3cf44809004149b

Observation 6127758c-8ad4-44f4-9215-885eb7169f91 · outbound

This paper cites Bag-of-visual-words and spatial extensions for land-use classification.

Context-Aware Multimodal Pretraining Bag-of-visual-words and spatial extensions for land-use classification

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.790291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.201604Z digest=sha256:0f10d6e6b14d575cb648fa04f6deaacaa352971011576f9e33de472424aae376

Observation 1b7a84d9-5cdd-41d1-9862-cf12abb957c2 · outbound

This paper cites TapNet: Neural network augmented with task-adaptive projection for few-shot learning.

Context-Aware Multimodal Pretraining TapNet: Neural network augmented with task-adaptive projection for few-shot learning

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.777850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.205556Z digest=sha256:92b75d7ad6f27a7a6970905c69ee932ff0512b4708c454720f3680c2593458c3

Observation df4d8c5c-6d8b-41b0-a767-f9c6521dc099 · outbound

This paper cites Sigmoid loss for language image pre-training.

Context-Aware Multimodal Pretraining Sigmoid loss for language image pre-training

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.765024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.210229Z digest=sha256:092b73c7ac0107a2eee6ef8abeff4eec97cffb03cc02408d41b421f7f5a27382

Observation e6204512-50f0-4d9c-b946-a3f74601fde8 · outbound

This paper cites Deepemd: Few-shot image classification with differen- tiable earth mover’s distance and structured classifiers.

Context-Aware Multimodal Pretraining Deepemd: Few-shot image classification with differen- tiable earth mover’s distance and structured classifiers

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.753232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:37:57.214117Z digest=sha256:aee68e3fda4ae4ff18f49c3b07c37b861d09ee693b1d259c2305efc7b11ced7e

Pith citing papers

Observation b9b8c456-4792-410d-b998-f07c992c7b0e · inbound

How to Merge Your Multimodal Models Over Time? cites this paper.

How to Merge Your Multimodal Models Over Time? Context-Aware Multimodal Pretraining

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:24:44.634180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T19:24:43.443062Z digest=sha256:6b1049f19ad14ea8cf2f3fdaffae304dd1c15f0686c164ec2162aaf65f6fc0ed