Pith. sign in

Paper Citation Record · LEDGER

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

As of 21 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2506.16673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16673 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:41:44.214299Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T03:31:33.964471Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T03:40:54.141854Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact3
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb0d8947-a5cd-476a-a08b-6a3b270f23f1 · outbound

This paper cites Layer Normalization.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:39.653398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:39.653398Z digest=sha256:1de3d1b3e1f7710f7077cfac8ab38403113e6ea3180b7d357adc9506073712d4

Observation 1210daea-a2a6-4b2a-8591-31dd5b9ec0f7 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Distilling the Knowledge in a Neural Network

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.603499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.603499Z digest=sha256:c6bd49742240a817155e1977027c646dc9c5ce2ccecc1a5ccec0ff51721a51f4

Observation 8d20089f-b83f-4d4d-af8c-0e25addf1014 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text su- pervision.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Scaling up visual and vision-language representation learning with noisy text su- pervision

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.787183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:40.818042Z digest=sha256:96844c461bc2cc5bb5f8151fa2bfe918d0c1121eed304176cdaf3c56c8ad42da

Observation e51020db-c6d2-4e42-8726-38decccecdf1 · outbound

This paper cites Learning multiple layers of features from tiny im- ages.Handbook of Systemic Autoimmune Diseases, 1(4),.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning multiple layers of features from tiny im- ages.Handbook of Systemic Autoimmune Diseases, 1(4),

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.777334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:40.919321Z digest=sha256:56cb090774a2978fac2419e50d44367d762335c3ac61fe9a0225e6e536844058

Observation 77274896-7a07-4c53-bc8b-be4c9ac55907 · outbound

This paper cites Clipath: Fine-tune clip with visual fea- ture fusion for pathology image analysis towards min- imizing data collection efforts.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clipath: Fine-tune clip with visual fea- ture fusion for pathology image analysis towards min- imizing data collection efforts

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.766897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:41.059755Z digest=sha256:56a5ab8435668ce7679329f3f88fff92ade14d7e871706f9ef18d4d57d3ecd88

Observation 9f089dd5-871e-45fd-8274-8df9a9fc5a1f · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.156877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.156877Z digest=sha256:40add5e53c2cc45f87fb159ad35d801bf030f3530cad79264d6536c620490403

Observation 942f3b57-1b3f-4ff5-aa75-1a36304224d2 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.756791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:41.226742Z digest=sha256:035768b3c3ccc6b1f03d6137eb60973fc2e79e920cbaafcbb7df2c9090afa5aa

Observation 34d6268d-ed76-42af-b3aa-77edaa48c71d · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.302093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.302093Z digest=sha256:3c88a78292354d93264c9d38badd262b4b8ce82cb6d8e0fb79703ad0a34db4f3

Observation 983c5e6a-388b-4d59-8928-9ce5afa2e772 · outbound

This paper cites Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.738963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:41.381057Z digest=sha256:2e5ba05f63bbfffc1830824820d5eec32219a8c03b1d1d5b19ab21c99456a5ad

Observation e4239e0a-3066-4fb3-9eaf-840639acf3d5 · outbound

This paper cites FoldGPT: Simple and Effective Large Language Model Compression Scheme.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge FoldGPT: Simple and Effective Large Language Model Compression Scheme

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.576024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.576024Z digest=sha256:636ea3050452e00da9786754ecba6fdf301c7de0a9e881d5ab6f868c077265fd

Observation 78e710c1-ceb4-43b7-a2bc-32a33ff584d7 · outbound

This paper cites Clip-branches: Interactive fine-tuning for text- image retrieval.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clip-branches: Interactive fine-tuning for text- image retrieval

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.721080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:41.680090Z digest=sha256:5a77618e57e05d4ef97c7e3dfe933559c4866d9cdb493e94eae34da38baeface

Observation 59700003-ea25-4573-9220-5fa415e5f49c · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ClipCap: CLIP Prefix for Image Captioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.828926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.828926Z digest=sha256:7c9a2d77ee6309ac248636487f2368bbd0985fcc90dbe1fb37041fa1ba8597d4

Observation 55ca4b34-9252-4ad3-9b9b-57979f93adda · outbound

This paper cites Compact language models via pruning and knowledge distillation.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Compact language models via pruning and knowledge distillation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.710512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:41.888375Z digest=sha256:fb047b6fd02031ce393d6d1b62e73b57c2797d86c88732752163aa4b41d67757

Observation 149bad71-fc54-4a54-929f-b7bfae66a016 · outbound

This paper cites CHiLS: Zero- shot image classification with hierarchical label sets.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CHiLS: Zero- shot image classification with hierarchical label sets

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.700607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:41.974012Z digest=sha256:922b4fe1b5a81613612010541ffed73d6b0093f6ee40d3136dbc6f25b76bf2b1

Observation d22b41c4-e8de-4f9d-b8d1-42281a2efdb2 · outbound

This paper cites an unresolved cited work.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:41:46.690119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:42.038682Z digest=sha256:f722b50f08ed97348222d0ce161ed0a13956d707418c3b118679c4c70fb15225

Observation 1c4ace28-b9e0-45ec-a7f1-0dbdb30aa7ca · outbound

This paper cites Language models are unsupervised multitask learners.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Language models are unsupervised multitask learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:42.417659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:42.417659Z digest=sha256:be65890b194a09cde6b4709a3f6746dd58b9df377b0f488f1fbef8632e06ae35

Observation e7913e60-7df6-4ba1-b8bb-b6570d741f7e · outbound

This paper cites Learning transferable visual models from nat- ural language supervision.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning transferable visual models from nat- ural language supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:42.525499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:42.525499Z digest=sha256:a77359e91dc8df63df98f84c1d5797632e3cae7a0409ef98c1ac1a5dbd9f1355

Observation a98d3e93-39aa-42b4-b253-5f8146c3135b · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.648692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:42.594229Z digest=sha256:89e6d399aca47f430a284cc4fbbdaa74309c416685f9a4fcdfea6080065d5f5d

Observation 87482803-e7a0-4381-952c-bae8b819c481 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.627340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:42.730440Z digest=sha256:a15228dab07a0e00ce2adf8bb8e357117289513848f10069f5bfd2f545ca854f

Observation f62f778a-ab77-42f5-ba9c-37c4ea823d8e · outbound

This paper cites Characterizing and avoid- ing negative transfer.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Characterizing and avoid- ing negative transfer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.597299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:42.923374Z digest=sha256:b0cdbb77671765591a89dab8fd49c6e9f102199d99b2da9685a6618368a2e4d6

Observation 8b136c55-2dc0-455d-87f1-4f9dd86c18f7 · outbound

This paper cites Learngene: From open-world to your learning task.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learngene: From open-world to your learning task

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.291286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:42.944359Z digest=sha256:e06ba72ffb0b96510de819a65d9751e922dee1dbe7aef34d4d0127c5ea6a38cd

Observation 8ec5d236-9419-43ea-ae21-f9ceb50cc882 · outbound

This paper cites Learngene: Inheriting Condensed Knowledge from the Ancestry Model to Descendant Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learngene: Inheriting Condensed Knowledge from the Ancestry Model to Descendant Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:43.035082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:43.035082Z digest=sha256:ddd05364025ca00b9bada26cc3ec45da694627adc71000a8b1c6b6a9d8b4503e

Observation a355136e-4ae3-4159-beae-2294fd60338a · outbound

This paper cites Vision transformers as probabilistic expan- sion from learngene.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Vision transformers as probabilistic expan- sion from learngene

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.158677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:43.198681Z digest=sha256:5593c26bed905279bc10e0dac9ebc506a6fdcca2fd1e65558be497ed63e5ed6a

Observation a023133f-f8b6-4034-acdd-d7b6fd96b9ca · outbound

This paper cites Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:43.321605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:43.321605Z digest=sha256:50650c9f7403acc6fb3115da670e68d438fa3b9c242d9980bb22da051634f031

Observation 996b9d26-f958-4d91-8ac7-493294c8f537 · outbound

This paper cites KIND: Knowledge Integration and Diversion for Training Decomposable Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge KIND: Knowledge Integration and Diversion for Training Decomposable Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:41:44.756302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:43.487929Z digest=sha256:3f8b49cddbddb9086bc6e9d16c018af5c1087a1ab9049a0ec64c0807cc8ead7c

Observation ceb347dc-1b90-4c89-aef2-57be930a8a3b · outbound

This paper cites CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:43.615143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:43.615143Z digest=sha256:754eeb503276cee54ab177c22efe71570af6f6c16e8e431b18f904c856d4e2ed

Observation 70dfdb34-7ca4-40d8-9c41-b1766c1acbcd · outbound

This paper cites an unresolved cited work.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:41:46.137224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:43.735490Z digest=sha256:ffec7b5d4fed7aa5aa8f6f2df05e173b8d3635ddcc2ec00071c97a6aef9547eb

Observation 2082631f-286b-4761-b5a2-68274aff7dbd · outbound

This paper cites Sigmoid loss for language image pre-training.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Sigmoid loss for language image pre-training

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:45.888512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:43.840052Z digest=sha256:734988568a85c74ed933f14066dcbe9085747c4792cc77b4e7ee8509a908c20e

Observation 13984ce2-e02d-4297-8d68-991175959fc3 · outbound

This paper cites Minivit: Compressing vision transformers with weight multiplexing.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Minivit: Compressing vision transformers with weight multiplexing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:45.576440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:43.907351Z digest=sha256:91fac2fc549dd975709907f21e7ebe76d15adbb7d84115620642dfb78e622c12

Observation db9e8f2c-82da-40c2-9525-65d048c539d0 · outbound

This paper cites CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:44.018505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:44.018505Z digest=sha256:52e954bf9d0b3798bbd48aaa16e687439473b500293bf23af165d389b8de074f

Observation 0ee55134-129c-450b-b7d0-5db5854a84ca · outbound

This paper cites Learning clip guided visual- text fusion transformer for video-based pedestrian attribute recognition.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning clip guided visual- text fusion transformer for video-based pedestrian attribute recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:45.302679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:44.116542Z digest=sha256:cb4b947520641ec47cd95e0b6fb6934f36acf985963e5684a673c5f2e4db3488

Observation 64496046-0093-4afe-b0ab-3ea71dd38fb8 · outbound

This paper cites 8 with the loss weightλ set to1.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge 8 with the loss weightλ set to1

Reference 47

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:41:44.546243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:44.214299Z digest=sha256:771bff3069cfe032989445542e7b9940207b18ccdcb81e680fe20185e9be534c

Observation 98ba2ed7-0ac0-4d72-8e15-208b5f67e433 · outbound

This paper cites Cats and dogs.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Cats and dogs

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:42.159844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:42.159844Z digest=sha256:f82ed46afaad9c90d6cd4a62bc9de8a98b139986d661db371c500efecbbb7d9b

Observation c309051c-b47c-4cb4-aa55-c012e17239d6 · outbound

This paper cites Microsoft coco: Com- mon objects in context.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Microsoft coco: Com- mon objects in context

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.496223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.496223Z digest=sha256:b5eefb474455973510679321a1d008c449299de6bcea6d49beb447a37ce1a53f

Observation a7a89bf2-2990-40f3-abc5-b4dcd4f3ff9b · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understand- ing.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Bert: Pre-training of deep bidirectional transformers for language understand- ing

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.809002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:40.172526Z digest=sha256:fe44230ef5d0ba3bc0412c5122e21cb01ddff8f756464e7d1199d7c513220f13

Observation 6005b785-7aaf-4522-895b-91e43ca15514 · outbound

This paper cites Clipping: Distilling clip-based models with a stu- dent base for video-language retrieval.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clipping: Distilling clip-based models with a stu- dent base for video-language retrieval

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.673029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:42.268837Z digest=sha256:8209acc67afda3eb2345c520d4abc3b718be9b986d637babbc419ef94255303c

Observation fefd7bed-e213-4290-b97c-7aa5e6336261 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.839841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:39.926867Z digest=sha256:c11df7ff7b950816e812b53db589daa73502766839a5ff83c98f7aee240d8113

Observation 02cce5a3-27e0-4c63-b4aa-0d7b7c05fe3e · outbound

This paper cites Openclip, July.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Openclip, July

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.798397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:40.705044Z digest=sha256:68a80a1f0fb2ba4b40de7c55a48c900407f5616d4c616c8c8043dd13feeab325

Observation 0f226c26-7bbe-4ff2-a102-1d8c3ee06c94 · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality- experts.Advances in Neural Information Processing Sys- tems, 35:32897–32912,.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Vlmo: Unified vision-language pre-training with mixture-of-modality- experts.Advances in Neural Information Processing Sys- tems, 35:32897–32912,

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.856955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:39.722053Z digest=sha256:928ade4e932a760a95b9ed5ea2fb37eb4cd71fc8ea4955d3dba345eca89fe723

Observation 8661aa16-e518-4425-bef4-13ac93d6980a · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Lawrence Zitnick, and Devi Parikh

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.616326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:42.843705Z digest=sha256:29d6b2d5033c9dd4beeb20462fb0be0555046c70de6d1b3a3034a45b4a323c96

Observation 0da41141-771f-40f8-aeda-fcc84f6ac6df · outbound

This paper cites Building variable- sized models via learngene pool.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Building variable- sized models via learngene pool

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.637519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:42.654304Z digest=sha256:143a4982b86a72952dbaa371d14548de76a7dbf3336fd6429d4f2f23b363b91e

Observation 73c67a20-8c1a-40ae-bdf5-6971a45b3686 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.275604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.275604Z digest=sha256:bf46598b50990e47bc5c54725519adf3e4398b8e8a3af1438d6862a317f358ac

Observation f24fbc33-3743-436b-89f8-99ac3074417b · outbound

This paper cites Transferring Core Knowledge via Learngenes.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Transferring Core Knowledge via Learngenes

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.379368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.379368Z digest=sha256:7bc1e2947c3e7280a4ee89cb3ecd2def02ad4f06db61a3fcf1636cbeab054ce8

Observation 59e47a0f-3e1a-43cd-b200-b7144e282dc5 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Reproducible scaling laws for contrastive language-image learning

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.829621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:40.005207Z digest=sha256:550b696a7e852fe94a4dad2459312fc60c9da8a07d0fc15981a21164ceeadb31

Observation 2a080e28-5759-4fe3-ad22-af9afa26175d · outbound

This paper cites Food-101–mining discriminative com- ponents with random forests.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Food-101–mining discriminative com- ponents with random forests

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:39.837013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:39.837013Z digest=sha256:884a86434025b68ae42e53341c629346e07f36155f28a0fcec60c06d3accf0d2

Observation 9612f030-0cc3-4c8e-b862-f57471afba35 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Imagenet: A large-scale hierarchical image database

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.089980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.089980Z digest=sha256:8587d866ac132f384703e91983c26283ff48b052f355ab21af4ec278221557f3

Observation 1645c199-146a-4897-910e-bac736c52528 · outbound

This paper cites WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:41:45.019297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:41:40.522210Z digest=sha256:851bdc6a48d0e77edd64532ca4439ec1408094b66b8f6e6364310da08804d976

Pith citing papers

Observation cddb2aaa-c23f-4a4c-ae77-8b645044def7 · inbound

Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions cites this paper.

Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:25:53.878606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T02:23:52.589354Z digest=sha256:7f0358ad491b86634e6fd8cdd13b356259953694817f6a30615db5a77bc52559

Observation ce73f04e-e77b-4e33-a86e-6f8d47c2ade2 · inbound

Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models cites this paper.

Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.143600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T03:31:33.964471Z digest=sha256:9a31149ad7b25f3ab333f62bd860969ce8c1c66bfb76d51bd2fb2bc0017c370b