Pith. sign in

Paper Citation Record · LEDGER

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2506.16673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16673 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:41:44.214299Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T03:31:33.964471Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T03:40:54.141854Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact3
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb0d8947-a5cd-476a-a08b-6a3b270f23f1 · outbound

This paper cites Layer Normalization.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:39.653398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:39.653398Z digest=sha256:ff936c22e3742c4e8b450499efb12a6500c07032e0df017e35fac610e39b9b83

Observation 1210daea-a2a6-4b2a-8591-31dd5b9ec0f7 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Distilling the Knowledge in a Neural Network

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.603499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.603499Z digest=sha256:26ebacdd241cfcffc5e7e5dca8ed04cf59f15e8f1349b6dad83e4cb82c272bd0

Observation 8d20089f-b83f-4d4d-af8c-0e25addf1014 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text su- pervision.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Scaling up visual and vision-language representation learning with noisy text su- pervision

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.787183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:40.818042Z digest=sha256:7caf23b2ee115f3d5646615941a415d986f97640473f3e1dee6541965d6462b4

Observation e51020db-c6d2-4e42-8726-38decccecdf1 · outbound

This paper cites Learning multiple layers of features from tiny im- ages.Handbook of Systemic Autoimmune Diseases, 1(4),.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning multiple layers of features from tiny im- ages.Handbook of Systemic Autoimmune Diseases, 1(4),

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.777334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:40.919321Z digest=sha256:e3136b7501e9750d09bb7953f6e205d2b160b0e37b0394f029fadc376e965bcd

Observation 77274896-7a07-4c53-bc8b-be4c9ac55907 · outbound

This paper cites Clipath: Fine-tune clip with visual fea- ture fusion for pathology image analysis towards min- imizing data collection efforts.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clipath: Fine-tune clip with visual fea- ture fusion for pathology image analysis towards min- imizing data collection efforts

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.766897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:41.059755Z digest=sha256:3aeed48a4ce5767d1c40af94541483d805ae88787067554ceee2fcccc59c1964

Observation 9f089dd5-871e-45fd-8274-8df9a9fc5a1f · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.156877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.156877Z digest=sha256:029cb9caa7312273d00b148830c2390ddd2b7a53d9aa15249a56d61c3e9c721e

Observation 942f3b57-1b3f-4ff5-aa75-1a36304224d2 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.756791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:41.226742Z digest=sha256:cb470d7455de15602ee25efa97446a5e9595cde4b347dc970ff4cd37fb730bdb

Observation 34d6268d-ed76-42af-b3aa-77edaa48c71d · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.302093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.302093Z digest=sha256:fd49b15de1dc8d9dc24a7fe3a7dfdd7ced78e21e6c6ad27993d1340d39893a95

Observation 983c5e6a-388b-4d59-8928-9ce5afa2e772 · outbound

This paper cites Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.738963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:41.381057Z digest=sha256:5a7d192433703730b3736e045031c5a5eb8428cc9f1e19786e368647d41905f0

Observation e4239e0a-3066-4fb3-9eaf-840639acf3d5 · outbound

This paper cites FoldGPT: Simple and Effective Large Language Model Compression Scheme.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge FoldGPT: Simple and Effective Large Language Model Compression Scheme

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.576024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.576024Z digest=sha256:1e6cc988be8223fbd5cdbf0aa7c12e921d5581ed264b5b3a0af9460e21a3c741

Observation 78e710c1-ceb4-43b7-a2bc-32a33ff584d7 · outbound

This paper cites Clip-branches: Interactive fine-tuning for text- image retrieval.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clip-branches: Interactive fine-tuning for text- image retrieval

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.721080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:41.680090Z digest=sha256:45b8f408e747938b646a5b822b901747ffc5e216352a568c6dc3871729a6bd40

Observation 59700003-ea25-4573-9220-5fa415e5f49c · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ClipCap: CLIP Prefix for Image Captioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.828926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.828926Z digest=sha256:fe626f5ad010e00def0b3c3e9604b94356460ae6adda6466326df97ce76c8a5b

Observation 55ca4b34-9252-4ad3-9b9b-57979f93adda · outbound

This paper cites Compact language models via pruning and knowledge distillation.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Compact language models via pruning and knowledge distillation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.710512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:41.888375Z digest=sha256:822d089165c1105aef60529817360cec9a2d07a31d7bcad4d983aa1eb1af33e4

Observation 149bad71-fc54-4a54-929f-b7bfae66a016 · outbound

This paper cites CHiLS: Zero- shot image classification with hierarchical label sets.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CHiLS: Zero- shot image classification with hierarchical label sets

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.700607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:41.974012Z digest=sha256:29aec1b114940a543117597b3be8d72b59bc2cf2742977b4be8fab717cfbb15a

Observation d22b41c4-e8de-4f9d-b8d1-42281a2efdb2 · outbound

This paper cites an unresolved cited work.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:41:46.690119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:42.038682Z digest=sha256:973f6470266b84adb5195a433aa13485fd4d12ccbdcdcf7bba022eb2b28e1175

Observation 1c4ace28-b9e0-45ec-a7f1-0dbdb30aa7ca · outbound

This paper cites Language models are unsupervised multitask learners.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Language models are unsupervised multitask learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:42.417659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:42.417659Z digest=sha256:6f4fac30522ac42e75da27dc55bb84e251cca8b5798099611a0293ef82275e1d

Observation e7913e60-7df6-4ba1-b8bb-b6570d741f7e · outbound

This paper cites Learning transferable visual models from nat- ural language supervision.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning transferable visual models from nat- ural language supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:42.525499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:42.525499Z digest=sha256:de96d9aa2170bd12486a8fc141135d0b60ec15b04f75070630d49b0b42f600a8

Observation a98d3e93-39aa-42b4-b253-5f8146c3135b · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.648692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:42.594229Z digest=sha256:8f05bde3fac256e0033c851a010e8c9b20edda30cbe7b8065a8cb34aeb4c8f78

Observation 87482803-e7a0-4381-952c-bae8b819c481 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.627340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:42.730440Z digest=sha256:50b655432508ac1d6c4459bdbdfdc3c3dcc13c41b177cb48b13e71cda00b3dd7

Observation f62f778a-ab77-42f5-ba9c-37c4ea823d8e · outbound

This paper cites Characterizing and avoid- ing negative transfer.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Characterizing and avoid- ing negative transfer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.597299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:42.923374Z digest=sha256:35e41d1542e49dbee9e6824a847455b9c7735099b21cc3ee21b335140203b2cd

Observation 8b136c55-2dc0-455d-87f1-4f9dd86c18f7 · outbound

This paper cites Learngene: From open-world to your learning task.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learngene: From open-world to your learning task

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.291286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:42.944359Z digest=sha256:5ab7fc7ca8512cb6c84c8fccb8db1700e8a27b8d721a2df682ecdd33c14120b8

Observation 8ec5d236-9419-43ea-ae21-f9ceb50cc882 · outbound

This paper cites Learngene: Inheriting Condensed Knowledge from the Ancestry Model to Descendant Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learngene: Inheriting Condensed Knowledge from the Ancestry Model to Descendant Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:43.035082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:43.035082Z digest=sha256:6ca3541b4b23e4dfba031740c5729ac26760565795151996e21e776105a67bd7

Observation a355136e-4ae3-4159-beae-2294fd60338a · outbound

This paper cites Vision transformers as probabilistic expan- sion from learngene.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Vision transformers as probabilistic expan- sion from learngene

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.158677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:43.198681Z digest=sha256:cec6d403d61d4037d1c9d2b67b5a0c33940e8cfae0ed9f9a3def3851c201556d

Observation a023133f-f8b6-4034-acdd-d7b6fd96b9ca · outbound

This paper cites Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:43.321605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:43.321605Z digest=sha256:83ac4557c6d78988d97f9a90c65448dc7a567253660ee36843fd91b44e29d758

Observation 996b9d26-f958-4d91-8ac7-493294c8f537 · outbound

This paper cites KIND: Knowledge Integration and Diversion for Training Decomposable Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge KIND: Knowledge Integration and Diversion for Training Decomposable Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:41:44.756302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:43.487929Z digest=sha256:cbae6cf564f9cf2b958a978d9441982f408eb33f76943ffd4bf5147ddebf020d

Observation ceb347dc-1b90-4c89-aef2-57be930a8a3b · outbound

This paper cites CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:43.615143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:43.615143Z digest=sha256:65c6076fb13ae4ea92645f369d4101ea97bc057cc03e8314d37b539fa83418a3

Observation 70dfdb34-7ca4-40d8-9c41-b1766c1acbcd · outbound

This paper cites an unresolved cited work.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:41:46.137224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:43.735490Z digest=sha256:ca8bea52c391386d044ffbdbeca8b88652ddcbd22ca13d97f2b3187bca598642

Observation 2082631f-286b-4761-b5a2-68274aff7dbd · outbound

This paper cites Sigmoid loss for language image pre-training.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Sigmoid loss for language image pre-training

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:45.888512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:43.840052Z digest=sha256:7b3ba823befe0009f8fe1dcd4c8ea388e9a1693664f34852b81aa6ac9b852f69

Observation 13984ce2-e02d-4297-8d68-991175959fc3 · outbound

This paper cites Minivit: Compressing vision transformers with weight multiplexing.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Minivit: Compressing vision transformers with weight multiplexing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:45.576440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:43.907351Z digest=sha256:2c16c54abbb10586bbee4fa488e6aa9c34ba94b320c8f745263eb1cbdba510f8

Observation db9e8f2c-82da-40c2-9525-65d048c539d0 · outbound

This paper cites CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:44.018505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:44.018505Z digest=sha256:badaba8cac4765443ba9f36e28b31aaaa5e190289129a6b3cb76bb9c5f874ab4

Observation 0ee55134-129c-450b-b7d0-5db5854a84ca · outbound

This paper cites Learning clip guided visual- text fusion transformer for video-based pedestrian attribute recognition.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning clip guided visual- text fusion transformer for video-based pedestrian attribute recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:45.302679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:44.116542Z digest=sha256:30e139d622a2a5d66404939fbb6ec7d37d1d75d094c06fcc1fab62348e2c0cf9

Observation 64496046-0093-4afe-b0ab-3ea71dd38fb8 · outbound

This paper cites 8 with the loss weightλ set to1.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge 8 with the loss weightλ set to1

Reference 47

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:41:44.546243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:44.214299Z digest=sha256:ec854fd2c2b52f9193464eca9163946e9ac2d8af819c4a21d666a0ec4a661a78

Observation 98ba2ed7-0ac0-4d72-8e15-208b5f67e433 · outbound

This paper cites Cats and dogs.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Cats and dogs

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:42.159844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:42.159844Z digest=sha256:81c917ccd3b78f0a5763115137357c07a69a23f5bec993ee8500b9f2dcb12bde

Observation c309051c-b47c-4cb4-aa55-c012e17239d6 · outbound

This paper cites Microsoft coco: Com- mon objects in context.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Microsoft coco: Com- mon objects in context

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.496223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.496223Z digest=sha256:470a3c9926289a6374b2c49165c2afc0eb83e99de99e0a383448bdd2820ee5f7

Observation a7a89bf2-2990-40f3-abc5-b4dcd4f3ff9b · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understand- ing.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Bert: Pre-training of deep bidirectional transformers for language understand- ing

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.809002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:40.172526Z digest=sha256:2ee16dfa3aec58ee3996f3ba3bb5cc10137cf24e7a0a5f1d27a037c987d57dbd

Observation 6005b785-7aaf-4522-895b-91e43ca15514 · outbound

This paper cites Clipping: Distilling clip-based models with a stu- dent base for video-language retrieval.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clipping: Distilling clip-based models with a stu- dent base for video-language retrieval

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.673029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:42.268837Z digest=sha256:84bb5b0ac45faacf5087d71da804860213842336de28f721f51d70b422b6a06e

Observation fefd7bed-e213-4290-b97c-7aa5e6336261 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.839841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:39.926867Z digest=sha256:8fd9569a1209e1271dde4776313b5418ca7622ce340c0cddc7bf1a9ea7b47e54

Observation 02cce5a3-27e0-4c63-b4aa-0d7b7c05fe3e · outbound

This paper cites Openclip, July.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Openclip, July

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.798397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:40.705044Z digest=sha256:64d3ec40f75019ec064c13c2614f7255650ee29aab4c63e097a7f46ac2c9649e

Observation 0f226c26-7bbe-4ff2-a102-1d8c3ee06c94 · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality- experts.Advances in Neural Information Processing Sys- tems, 35:32897–32912,.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Vlmo: Unified vision-language pre-training with mixture-of-modality- experts.Advances in Neural Information Processing Sys- tems, 35:32897–32912,

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.856955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:39.722053Z digest=sha256:9769658f763070639909849abe02a9fc7dc981b3d708aad9f785823133c27a61

Observation 8661aa16-e518-4425-bef4-13ac93d6980a · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Lawrence Zitnick, and Devi Parikh

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.616326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:42.843705Z digest=sha256:1feebdb95f1e584c477e4c4085463dc316d3fbb2222736f8fd9aea0e73d726bf

Observation 0da41141-771f-40f8-aeda-fcc84f6ac6df · outbound

This paper cites Building variable- sized models via learngene pool.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Building variable- sized models via learngene pool

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.637519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:42.654304Z digest=sha256:4ec1f62ff1beb2ab491f68202bf5f239d47b9a97fe39d4f6d99ea26c622e70a2

Observation 73c67a20-8c1a-40ae-bdf5-6971a45b3686 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.275604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.275604Z digest=sha256:40a6793c2cd9af514d6f176537f74d29af474d9f3f65774dc673cbfeabaeea48

Observation f24fbc33-3743-436b-89f8-99ac3074417b · outbound

This paper cites Transferring Core Knowledge via Learngenes.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Transferring Core Knowledge via Learngenes

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.379368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.379368Z digest=sha256:cc65e3dea2f8238da7d9a3ad08be41c9db525dc66feb7eb7885fc13a51d33d89

Observation 59e47a0f-3e1a-43cd-b200-b7144e282dc5 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Reproducible scaling laws for contrastive language-image learning

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:41:46.829621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:40.005207Z digest=sha256:b253ad6ec3e5881328654e26cdab1097b9e5fd64b5f69d327772b780f1f6a6c8

Observation 2a080e28-5759-4fe3-ad22-af9afa26175d · outbound

This paper cites Food-101–mining discriminative com- ponents with random forests.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Food-101–mining discriminative com- ponents with random forests

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:39.837013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:39.837013Z digest=sha256:d71d4e05ce10ef269c8a6bbd688561ca62797cdcfb85c0e5a3c6f09c33628a1b

Observation 9612f030-0cc3-4c8e-b862-f57471afba35 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Imagenet: A large-scale hierarchical image database

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:40.089980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:40.089980Z digest=sha256:0043cad674c71349007fa90ef15fc5c9202659f80a56f81db8e5c3620e612230

Observation 1645c199-146a-4897-910e-bac736c52528 · outbound

This paper cites WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:41:45.019297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:41:40.522210Z digest=sha256:8fcd6921cec08be497efe6b62bf45452543297133a3177f2ec5287ce7292ce12

Pith citing papers

Observation cddb2aaa-c23f-4a4c-ae77-8b645044def7 · inbound

Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions cites this paper.

Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:25:53.878606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T02:23:52.589354Z digest=sha256:66d3b3e805c44ed25519432cd5903eb35cbd075d4fefee3b30907748ea91f789

Observation ce73f04e-e77b-4e33-a86e-6f8d47c2ade2 · inbound

Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models cites this paper.

Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.143600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T03:31:33.964471Z digest=sha256:1407476b87758738b56b3027c8eaa72a568fcaecfb4b787fa7154e454f14b96b