Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:41:44.214299Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2506.16673.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:41:44.214299Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-11T03:31:33.964471Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T03:40:54.141854Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bb0d8947-a5cd-476a-a08b-6a3b270f23f1 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Layer Normalization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1210daea-a2a6-4b2a-8591-31dd5b9ec0f7 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Distilling the Knowledge in a Neural Network
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d20089f-b83f-4d4d-af8c-0e25addf1014 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Scaling up visual and vision-language representation learning with noisy text su- pervision
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e51020db-c6d2-4e42-8726-38decccecdf1 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning multiple layers of features from tiny im- ages.Handbook of Systemic Autoimmune Diseases, 1(4),
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 77274896-7a07-4c53-bc8b-be4c9ac55907 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clipath: Fine-tune clip with visual fea- ture fusion for pathology image analysis towards min- imizing data collection efforts
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f089dd5-871e-45fd-8274-8df9a9fc5a1f · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 942f3b57-1b3f-4ff5-aa75-1a36304224d2 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34d6268d-ed76-42af-b3aa-77edaa48c71d · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 983c5e6a-388b-4d59-8928-9ce5afa2e772 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e4239e0a-3066-4fb3-9eaf-840639acf3d5 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge FoldGPT: Simple and Effective Large Language Model Compression Scheme
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78e710c1-ceb4-43b7-a2bc-32a33ff584d7 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clip-branches: Interactive fine-tuning for text- image retrieval
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59700003-ea25-4573-9220-5fa415e5f49c · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ClipCap: CLIP Prefix for Image Captioning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ca4b34-9252-4ad3-9b9b-57979f93adda · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Compact language models via pruning and knowledge distillation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 149bad71-fc54-4a54-929f-b7bfae66a016 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CHiLS: Zero- shot image classification with hierarchical label sets
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d22b41c4-e8de-4f9d-b8d1-42281a2efdb2 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1c4ace28-b9e0-45ec-a7f1-0dbdb30aa7ca · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Language models are unsupervised multitask learners
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7913e60-7df6-4ba1-b8bb-b6570d741f7e · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning transferable visual models from nat- ural language supervision
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a98d3e93-39aa-42b4-b253-5f8146c3135b · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87482803-e7a0-4381-952c-bae8b819c481 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f62f778a-ab77-42f5-ba9c-37c4ea823d8e · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Characterizing and avoid- ing negative transfer
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8b136c55-2dc0-455d-87f1-4f9dd86c18f7 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learngene: From open-world to your learning task
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ec5d236-9419-43ea-ae21-f9ceb50cc882 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learngene: Inheriting Condensed Knowledge from the Ancestry Model to Descendant Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a355136e-4ae3-4159-beae-2294fd60338a · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Vision transformers as probabilistic expan- sion from learngene
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a023133f-f8b6-4034-acdd-d7b6fd96b9ca · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 996b9d26-f958-4d91-8ac7-493294c8f537 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge KIND: Knowledge Integration and Diversion for Training Decomposable Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ceb347dc-1b90-4c89-aef2-57be930a8a3b · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70dfdb34-7ca4-40d8-9c41-b1766c1acbcd · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2082631f-286b-4761-b5a2-68274aff7dbd · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Sigmoid loss for language image pre-training
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13984ce2-e02d-4297-8d68-991175959fc3 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Minivit: Compressing vision transformers with weight multiplexing
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db9e8f2c-82da-40c2-9525-65d048c539d0 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee55134-129c-450b-b7d0-5db5854a84ca · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Learning clip guided visual- text fusion transformer for video-based pedestrian attribute recognition
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 64496046-0093-4afe-b0ab-3ea71dd38fb8 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge 8 with the loss weightλ set to1
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98ba2ed7-0ac0-4d72-8e15-208b5f67e433 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Cats and dogs
Reference 2001
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c309051c-b47c-4cb4-aa55-c012e17239d6 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Microsoft coco: Com- mon objects in context
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a89bf2-2990-40f3-abc5-b4dcd4f3ff9b · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Bert: Pre-training of deep bidirectional transformers for language understand- ing
Reference 2009
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6005b785-7aaf-4522-895b-91e43ca15514 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Clipping: Distilling clip-based models with a stu- dent base for video-language retrieval
Reference 2012
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fefd7bed-e213-4290-b97c-7aa5e6336261 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 2014
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02cce5a3-27e0-4c63-b4aa-0d7b7c05fe3e · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Openclip, July
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0f226c26-7bbe-4ff2-a102-1d8c3ee06c94 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Vlmo: Unified vision-language pre-training with mixture-of-modality- experts.Advances in Neural Information Processing Sys- tems, 35:32897–32912,
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8661aa16-e518-4425-bef4-13ac93d6980a · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Lawrence Zitnick, and Devi Parikh
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0da41141-771f-40f8-aeda-fcc84f6ac6df · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Building variable- sized models via learngene pool
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 73c67a20-8c1a-40ae-bdf5-6971a45b3686 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f24fbc33-3743-436b-89f8-99ac3074417b · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Transferring Core Knowledge via Learngenes
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59e47a0f-3e1a-43cd-b200-b7144e282dc5 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Reproducible scaling laws for contrastive language-image learning
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a080e28-5759-4fe3-ad22-af9afa26175d · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Food-101–mining discriminative com- ponents with random forests
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9612f030-0cc3-4c8e-b862-f57471afba35 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge Imagenet: A large-scale hierarchical image database
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1645c199-146a-4897-910e-bac736c52528 · outbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cddb2aaa-c23f-4a4c-ae77-8b645044def7 · inbound
Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce73f04e-e77b-4e33-a86e-6f8d47c2ade2 · inbound
Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.