Pith. sign in

Paper Citation Record · LEDGER

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers

As of 13 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2411.14789.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14789 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:59:52.281615Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3667f728-116f-4d70-8df0-b86563989a6b · outbound

This paper cites Learning transferable visual models from natural language supervision.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.114205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.114205Z digest=sha256:f9ba3343abd214cf9b4c7476b33ab6984076e7b785f75cd22180dc0adc3537bf

Observation 7f3b4057-6abf-4e5c-ba0f-6635cd45248c · outbound

This paper cites Mobileclip: Fast image-text models through multi-modal reinforced training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Mobileclip: Fast image-text models through multi-modal reinforced training

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.788275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.119201Z digest=sha256:983ad4667d746c4c6c08a9305061b4c6e396f671c611fa3b7e8372d95ea51e1a

Observation bd51ce5b-e1d3-47b9-939d-baad2c37c84a · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.123160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.123160Z digest=sha256:ccfbf9e151067dffbdf3a0ba365ed81f4e9e3bd6f40d450ca5237eb039b688cf

Observation 95ba7f40-9837-4a6a-96f4-48f3644cb71e · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.127302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.127302Z digest=sha256:a6014122969b89a4953fbfad5e9959f9b773bbad4b93b95b2f146be22bc3eed6

Observation 772f5775-3f49-4cc9-96ca-86ee78d549a0 · outbound

This paper cites Slip: Self-supervision meets language-image pre-training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Slip: Self-supervision meets language-image pre-training

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.768761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.131546Z digest=sha256:9bc50b8a33511ea602ea4d9a1f7f6f9314aa66b63065fda0b528c53337aeb8f4

Observation 9e9d7ef6-4019-4fdb-ad6a-63b163c52140 · outbound

This paper cites Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.135248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.135248Z digest=sha256:85c54e288ac323046531e1104e53ef4108691fe6cfe3b621e2d4a236d14bd63a

Observation deefb527-540a-4655-a01e-204eb4de7628 · outbound

This paper cites Unified contrastive learning in image-text-label space.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Unified contrastive learning in image-text-label space

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.757540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.139565Z digest=sha256:3985d3b8fba64accf399f77a7831db236e080316d6ee0772155d8a8055c25a5c

Observation befa4a88-7088-4800-977f-9293a0e45170 · outbound

This paper cites Sigmoid loss for language image pre- training.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Sigmoid loss for language image pre- training

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.745608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.142687Z digest=sha256:eb7d7ff402a67b59df383e0737f7315ac49ff9bc3b314ba2bbbbaacb102b70f4

Observation d224ba9a-6194-4892-b457-f70d39b92c84 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.733608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.145599Z digest=sha256:47521ab594e200ba2b826e334790870a39da4b1aa7d1fc401a5688e753c8e76d

Observation d2113434-c7b6-416a-b61d-d75c170e37ad · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Distilling the Knowledge in a Neural Network

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.148493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.148493Z digest=sha256:0f0c2df958f1c22b432c0f1a27c151dafaf1377e34a1b5277fcee2261b18fc7e

Observation 0e814fee-c16d-467d-a4cc-56dc16025f96 · outbound

This paper cites Tinyclip: Clip distillation via affinity mimicking and weight inheritance.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Tinyclip: Clip distillation via affinity mimicking and weight inheritance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.151714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.151714Z digest=sha256:775d3b7ae3e65b8f975b878382f036e9b78dee0425c396f4ea82366b23b7ce9d

Observation d09a2960-9757-4414-b699-6877ec18ba32 · outbound

This paper cites Clip-kd: An empirical study of clip model distillation.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Clip-kd: An empirical study of clip model distillation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.154940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.154940Z digest=sha256:d40651be2de86897314af9270f9c8acbdac25ac09012eef991050e46210225ca

Observation 06934296-ab32-4eb8-921d-1b70242d165a · outbound

This paper cites Data Filtering Networks.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Data Filtering Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.157799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.157799Z digest=sha256:966c6850fce217644087a08be1118716ebaf17ad994094b4a521bcf1007dda51

Observation 6da76b83-de18-419b-b00f-2480161d7073 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Datacomp: In search of the next generation of multimodal datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.161256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.161256Z digest=sha256:6a17cff02f45bf438122274fcb090a7771d2e0627ccc9b0646d0f9a6454f009a

Observation 0046f2d9-6ffa-40b8-8898-93617acc3956 · outbound

This paper cites Alip: Adaptive language-image pre-training with synthetic caption.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Alip: Adaptive language-image pre-training with synthetic caption

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.689960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.164341Z digest=sha256:2c9a33e491d8ec3a334482f2994dbdf06fe8235e6b124c596919af7f489abd3d

Observation 99f775da-c713-4249-abd7-54cd86e75a8d · outbound

This paper cites VeCLIP: Improving CLIP Training via Visual-enriched Captions.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers VeCLIP: Improving CLIP Training via Visual-enriched Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.167679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.167679Z digest=sha256:821cf05b420bebe8391c4bd8a04d15bee7875cb92a2a2240aa0703dcc9143b97

Observation 36e15436-1244-43c3-8391-60b30af1edca · outbound

This paper cites Metaformer is actually what you need for vision.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Metaformer is actually what you need for vision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.679633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.171120Z digest=sha256:328af7e0419557db75254555758e03ffac5affff91096ef2cfd809cb268a8795

Observation c8118bb7-c484-474a-83da-f2b90090f27f · outbound

This paper cites Deeply supervised salient object detection with short connections.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Deeply supervised salient object detection with short connections

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.668728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.174562Z digest=sha256:90e8047a9a34249dbf601ad31ded81117b227af85f0ba0a39a6a25379944386d

Observation ecfb34f2-c296-4364-86e2-a67e3bae9621 · outbound

This paper cites Cascaded partial decoder for fast and accurate salient object detection.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Cascaded partial decoder for fast and accurate salient object detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.657079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.178714Z digest=sha256:259ea809f40559544a6120cba9b781717e6410e21b0112d380980a709626dad9

Observation ec24a1d5-3abc-4585-ae60-f988163ab6d0 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.182930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.182930Z digest=sha256:f60d044dcdc069fcf91d6e9bebf51cf2210aa51d63623912d4eaeded45f02f9d

Observation fd94dcaf-8fd9-41bc-acb9-11a1dab9fa50 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.186821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.186821Z digest=sha256:a1caebb50377206507b6eac71e48c34d9029e17afafba03e51db6bbf1a773fbe

Observation abbf80c6-af9b-4da1-b80a-8f4a4d48a0d3 · outbound

This paper cites Less is more: Pay less attention in vision transformers.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Less is more: Pay less attention in vision transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.637839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.190581Z digest=sha256:eb0b4e606856bcc225157406fccc5afe6d7be5ab2858799dd7a5f940f0b9c2db

Observation fd86234b-74ed-4a34-821a-ad2ef9365662 · outbound

This paper cites Cmt: Convolutional neural networks meet vision transformers.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Cmt: Convolutional neural networks meet vision transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.625709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.194481Z digest=sha256:89f8f2bf05c5e09a6f94a4da6eba6f4a77bbf7e1bbd1399327f9bfd2d5453ed9

Observation ddca9952-036c-4d8c-8fc8-ac33cd6bcebb · outbound

This paper cites Fastvit: A fast hybrid vision transformer using structural reparameterization.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Fastvit: A fast hybrid vision transformer using structural reparameterization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.613709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.200277Z digest=sha256:0badde48061df4506f3a3ea95820412ee6b11a88f9a7049a4e897bc8fc7a40d2

Observation 97a4c6e0-62c2-4894-9c15-4ad64edd5671 · outbound

This paper cites Universal Transformers.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Universal Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.205965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.205965Z digest=sha256:23cc39fee7503ed529c406da15f5fad130585300b113c81a7e65b55c53dd651e

Observation e5f31a3b-c752-4246-8d4a-651867b554eb · outbound

This paper cites Perceiver: General perception with iterative attention.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Perceiver: General perception with iterative attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.210190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.210190Z digest=sha256:2e2ceeface2cbe93a8431b102ec1592632f8986428d9b62c58470603ee27d0c4

Observation 645b9ae5-35a5-4397-b4ef-4028d0cb4d43 · outbound

This paper cites Sharing low rank conformer weights for tiny always-on ambient speech recognition models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Sharing low rank conformer weights for tiny always-on ambient speech recognition models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.593450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.214395Z digest=sha256:25368abb680732dd95ba46828fa7c9c70ff97d63267293725b80ee9068f93552

Observation de702da3-e804-42f5-9342-b6696ed0b4fb · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.219465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.219465Z digest=sha256:331276ddf23ba3f6bfc04176621babbdddc70a6cd8cb983d6b188e96238fdd3e

Observation 19bc353d-c96a-4333-9622-b5e048a40494 · outbound

This paper cites Simplifying transformer blocks.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Simplifying transformer blocks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.581833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.223873Z digest=sha256:f95148e8b6eb6b6fdf3fc88e8c557306f856d93df6d604d735b6ffe39e290218

Observation 26350de0-be99-4fa9-aadd-01befde4bec1 · outbound

This paper cites Attention is all you need.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Attention is all you need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.227589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.227589Z digest=sha256:519b924da4e2c5e0d7024d5dccde3e0d291611f797ea3e92ccb76d8579082ce0

Observation 011605a2-487f-4b23-815c-7495e8f36552 · outbound

This paper cites The shaped transformer: Attention models in the infinite depth-and-width limit.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers The shaped transformer: Attention models in the infinite depth-and-width limit

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.564010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.232264Z digest=sha256:97bf37d22a3db765e07a33abd2df7e6a7333315c5884ded305bdbb98615d4a1c

Observation 5d5940fc-69ea-4bae-852d-4586dd2e2c14 · outbound

This paper cites Augmenting Self-attention with Persistent Memory.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Augmenting Self-attention with Persistent Memory

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.236035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.236035Z digest=sha256:aed82f13d8de1b4724b959c45b5220ee565a019bd1281e7d867a2fdd167ed6f0

Observation d4744f33-043d-4ad0-b297-e92d696d6a8e · outbound

This paper cites Beyond Scaling Laws: Understanding Transformer Performance with Associative Memory.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Beyond Scaling Laws: Understanding Transformer Performance with Associative Memory

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.240108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.240108Z digest=sha256:a59a24209db6d6b371c7eeb7a8cef927ba683bc69f61c9fc03dc9a930e18443b

Observation 4a2d84b0-0353-4449-84eb-a5c79032403e · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers LoRA: Low-Rank Adaptation of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.244204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.244204Z digest=sha256:92c751ede0081627886bf7b1660d541842e7d6a4c7651f9096a6675dfd7df997

Observation a43af1cc-e3a0-4340-9a0b-9d5d8037fd69 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.247115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.247115Z digest=sha256:71f5cc1e46be7a9a75c25a6d50a877b3ca39f9f8aa9d68eebd7b67fdb6053d3e

Observation 57d172d9-f745-4b17-acf7-c7d30bbe706b · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.250581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.250581Z digest=sha256:f057910b73454039730e237b28ca4fcbc609e6aa722a7f923934ddbd673de622

Observation 1ed292a2-cff2-4386-9a2c-90537d6deea3 · outbound

This paper cites Rils: Masked visual reconstruction in language semantic space.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Rils: Masked visual reconstruction in language semantic space

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.543998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.253198Z digest=sha256:de01df1d1945d6c57d11e3dd0a126658121e3c5de57533e9046ce4b7a0cdf47c

Observation 8e900ed7-dc7f-4e8d-b091-74c8b1ba186c · outbound

This paper cites Extract free dense labels from clip.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Extract free dense labels from clip

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.255907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.255907Z digest=sha256:719b43d11851542c6a959fad4de8c30dd2699dfd970b8a5e0c3ad5c6ca97ab0b

Observation a2a50556-d133-47dc-bc8d-39e181098183 · outbound

This paper cites Openclip repository, 2021.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Openclip repository, 2021

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.520710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.258639Z digest=sha256:19c44bfe1493fd2c051afb4be457b728631139a6a20081c5c22eb09117ce6f54

Observation f7f7c6fd-a8e6-4aab-911b-4e83350e8020 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Imagenet: A large-scale hierarchical image database

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.261429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.261429Z digest=sha256:1ff226feb75f74e86253d6ef87f03da99dfcd599384db4f278493131309c2394

Observation 4241ff8b-9f35-40f9-bf3d-38bb0188fd6d · outbound

This paper cites Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.264271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.264271Z digest=sha256:221d221c11ae555687f4874096f6aff55182e0bb39a45f3ed305dbbbd1051d3d

Observation f0d35b14-1eaa-45ee-b25b-470128a85c44 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers The many faces of robustness: A critical analysis of out-of-distribution generalization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.267287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.267287Z digest=sha256:b6d368cce77138c775cc3fb38c6aef478bb9697192cfbf76164074ffc8510fc2

Observation c38f53c5-6c36-4c81-917b-9db95a4c2a22 · outbound

This paper cites Learning robust global representations by penalizing local predictive power.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Learning robust global representations by penalizing local predictive power

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.270399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.270399Z digest=sha256:6505da36da2aecaa9b35620984d40dba34ff3c77af39093d37c7e75cafcc5606

Observation 7d573126-e6cb-48ac-b674-e921a159f385 · outbound

This paper cites Microsoft coco: Common objects in context.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Microsoft coco: Common objects in context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.273910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.273910Z digest=sha256:030478e53a2d423833e4a719988bb5d2de145fb95a796166b3a49e4ed2a48f50

Observation 84f5fab0-faaf-4a37-ba18-c40afd5f6426 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.277763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.277763Z digest=sha256:6eee72def396f3adb1a6403e600d68c32edf29abf144e9974f65ba328dacd5b4

Observation d1580e7f-5c9a-4e48-959a-6db2de7f8d86 · outbound

This paper cites Randaugment: Practical automated data augmentation with a reduced search space.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Randaugment: Practical automated data augmentation with a reduced search space

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:59:52.452649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:59:52.281615Z digest=sha256:085c6763e411c65feb6461270cfae03bc54a88b51cf768a17972de614435b3ba

Pith citing papers

No inbound Pith citation observations are available.