Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:59:52.281615Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2411.14789.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:59:52.281615Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3667f728-116f-4d70-8df0-b86563989a6b · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Learning transferable visual models from natural language supervision
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f3b4057-6abf-4e5c-ba0f-6635cd45248c · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Mobileclip: Fast image-text models through multi-modal reinforced training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bd51ce5b-e1d3-47b9-939d-baad2c37c84a · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95ba7f40-9837-4a6a-96f4-48f3644cb71e · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 772f5775-3f49-4cc9-96ca-86ee78d549a0 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Slip: Self-supervision meets language-image pre-training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9e9d7ef6-4019-4fdb-ad6a-63b163c52140 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deefb527-540a-4655-a01e-204eb4de7628 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Unified contrastive learning in image-text-label space
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation befa4a88-7088-4800-977f-9293a0e45170 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Sigmoid loss for language image pre- training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d224ba9a-6194-4892-b457-f70d39b92c84 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2113434-c7b6-416a-b61d-d75c170e37ad · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Distilling the Knowledge in a Neural Network
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e814fee-c16d-467d-a4cc-56dc16025f96 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Tinyclip: Clip distillation via affinity mimicking and weight inheritance
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d09a2960-9757-4414-b699-6877ec18ba32 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Clip-kd: An empirical study of clip model distillation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06934296-ab32-4eb8-921d-1b70242d165a · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Data Filtering Networks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da76b83-de18-419b-b00f-2480161d7073 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Datacomp: In search of the next generation of multimodal datasets
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0046f2d9-6ffa-40b8-8898-93617acc3956 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Alip: Adaptive language-image pre-training with synthetic caption
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 99f775da-c713-4249-abd7-54cd86e75a8d · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers VeCLIP: Improving CLIP Training via Visual-enriched Captions
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36e15436-1244-43c3-8391-60b30af1edca · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Metaformer is actually what you need for vision
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8118bb7-c484-474a-83da-f2b90090f27f · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Deeply supervised salient object detection with short connections
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ecfb34f2-c296-4364-86e2-a67e3bae9621 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Cascaded partial decoder for fast and accurate salient object detection
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ec24a1d5-3abc-4585-ae60-f988163ab6d0 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd94dcaf-8fd9-41bc-acb9-11a1dab9fa50 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Swin transformer: Hierarchical vision transformer using shifted windows
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abbf80c6-af9b-4da1-b80a-8f4a4d48a0d3 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Less is more: Pay less attention in vision transformers
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fd86234b-74ed-4a34-821a-ad2ef9365662 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Cmt: Convolutional neural networks meet vision transformers
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ddca9952-036c-4d8c-8fc8-ac33cd6bcebb · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Fastvit: A fast hybrid vision transformer using structural reparameterization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 97a4c6e0-62c2-4894-9c15-4ad64edd5671 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Universal Transformers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5f31a3b-c752-4246-8d4a-651867b554eb · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Perceiver: General perception with iterative attention
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645b9ae5-35a5-4397-b4ef-4028d0cb4d43 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Sharing low rank conformer weights for tiny always-on ambient speech recognition models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation de702da3-e804-42f5-9342-b6696ed0b4fb · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Conformer: Convolution-augmented Transformer for Speech Recognition
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19bc353d-c96a-4333-9622-b5e048a40494 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Simplifying transformer blocks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 26350de0-be99-4fa9-aadd-01befde4bec1 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Attention is all you need
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 011605a2-487f-4b23-815c-7495e8f36552 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers The shaped transformer: Attention models in the infinite depth-and-width limit
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5d5940fc-69ea-4bae-852d-4586dd2e2c14 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Augmenting Self-attention with Persistent Memory
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4744f33-043d-4ad0-b297-e92d696d6a8e · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Beyond Scaling Laws: Understanding Transformer Performance with Associative Memory
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2d84b0-0353-4449-84eb-a5c79032403e · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers LoRA: Low-Rank Adaptation of Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a43af1cc-e3a0-4340-9a0b-9d5d8037fd69 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57d172d9-f745-4b17-acf7-c7d30bbe706b · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ed292a2-cff2-4386-9a2c-90537d6deea3 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Rils: Masked visual reconstruction in language semantic space
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8e900ed7-dc7f-4e8d-b091-74c8b1ba186c · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Extract free dense labels from clip
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a50556-d133-47dc-bc8d-39e181098183 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Openclip repository, 2021
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f7f7c6fd-a8e6-4aab-911b-4e83350e8020 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Imagenet: A large-scale hierarchical image database
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4241ff8b-9f35-40f9-bf3d-38bb0188fd6d · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d35b14-1eaa-45ee-b25b-470128a85c44 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers The many faces of robustness: A critical analysis of out-of-distribution generalization
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c38f53c5-6c36-4c81-917b-9db95a4c2a22 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Learning robust global representations by penalizing local predictive power
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d573126-e6cb-48ac-b674-e921a159f385 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Microsoft coco: Common objects in context
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f5fab0-faaf-4a37-ba18-c40afd5f6426 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1580e7f-5c9a-4e48-959a-6db2de7f8d86 · outbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Randaugment: Practical automated data augmentation with a reduced search space
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.